← 开源
Nutlope

llama-ocr

Document to Markdown OCR library with Llama 3.2 vision

AI EngineeringGive agent knowledgeTypeScript
在 GitHub 打开
增长势头
+024 小时新增 Star0.0%
2.44k
Star
234
Fork
+1
本周
3
贡献者
创建于 2024-11-12 · 更新于 2026-10-05 · 今日第 8491 名
主要开发者
README

Llama OCR

An npm library to run OCR for free with Llama 3.2 Vision.

Current version


Installation

npm i llama-ocr

Usage

import { ocr } from "llama-ocr";

const markdown = await ocr({
  filePath: "./trader-joes-receipt.jpg", // path to your image (soon PDF!)
  apiKey: process.env.TOGETHER_API_KEY, // Together AI API key
});

Hosted Demo

We have a hosted demo at LlamaOCR.com where you can try it out!

How it works

This library uses the free Llama 3.2 endpoint from Together AI to parse images and return markdown. Paid endpoints for Llama 3.2 11B and Llama 3.2 90B are also available for faster performance and higher rate limits.

You can control this with the model option which is set to Llama-3.2-90B-Vision by default but can also accept free or Llama-3.2-11B-Vision.

Roadmap

  • [x] Add support for local images OCR
  • [x] Add support for remote images OCR
  • [ ] Add support for single page PDFs
  • [ ] Add support for multi-page PDFs OCR (take screenshots of PDF & feed to vision model)
  • [ ] Add support for JSON output in addition to markdown

Credit

This project was inspired by Zerox. Go check them out!