Preferences

Hijacking this thread: what's currently the cheapest way to get structured data out of a PDF?

I assume there's some reasonable tool out there to convert PDFs to Markup and than feed it to some LLM API with okay costs (Gemini? DeepSeek?). Any suggestions?


https://mistral.ai/news/mistral-ocr , recent release. Its been a step function improvement for my pipelines
I’m feeding pdfs directly to Gemini to extract tables and so far the results are pretty good. There was a post on HN a few days ago about using Gemini for this task.

This item has no comments currently.

Keyboard Shortcuts

Story Lists

j
Next story
k
Previous story
Shift+j
Last story
Shift+k
First story
o Enter
Go to story URL
c
Go to comments
u
Go to author

Navigation

Shift+t
Go to top stories
Shift+n
Go to new stories
Shift+b
Go to best stories
Shift+a
Go to Ask HN
Shift+s
Go to Show HN

Miscellaneous

?
Show this modal