Schedule
3:00 pm
The Life-Changing Magic of LLM Document Extraction
Does this PDF spark joy? For most companies, the answer is no. They have piles of documents that people still read by hand: contracts, invoices, forms, old scans. With an LLM and a Pydantic model, you can turn them into clean data in an afternoon. Then you run the demo again, and the answer changes. Look closer, and some values are not on the page at all: the model made them up to fill your schema.
This talk follows one document from that first joyful demo to an extraction pipeline that holds up in production. We look at why temperature 0 does not make an LLM deterministic, and how a strict schema can push a model to invent data.
Then we tidy up the way we tidy any software: break it into the smallest units. Use code where you can and the LLM only where you must. Extract one field at a time, allow "not stated", and give every field a link to where it came from. Keep what sparks joy, and validate the rest.
Host
Juanette Evans
Business Development Specialist
Xebia
Guests

Andy Ho
AI Engineer
Xebia