Bachelor's or master's thesis: document AI for invoices (m/f/d)
Delivery notes, invoices, waybills: use your thesis to find out how reliably AI can read business documents – and where it reaches its limits.
What it’s about
A crooked scan of a delivery note, an invoice with three pages of line items, a waybill with a handwritten remark: in many businesses, people still type documents like these into the ERP system field by field. Modern language and vision models can take that over – but how accurately, at what scan quality, and with how much checking afterwards? That is the question behind this thesis.
vona is a web agency from Leipzig working remote-first for small and mid-sized companies, established businesses, trades and startups across the German-speaking countries. We combine AI integration, development, design and SEO in one team, and we give honest advice – even when the answer is “this isn’t worth it”.
For a logistics provider, we automated document processing: up to 400 delivery notes, waybills and customs documents a day run through OCR and a locally hosted language model. The result: 85% less manual data entry, 12 seconds instead of 4 minutes per document, and 99.2% extraction accuracy. Your work can build on this. One possible research question: how do classic OCR pipelines followed by a language model compare with multimodal models that read the image directly – in terms of accuracy, cost and the kinds of errors they make? Your own ideas, for instance on table recognition or on models assessing their own confidence, are welcome.
Your tasks
- Put together an anonymized test corpus covering different document types and scan qualities, and define which fields need to be extracted.
- Build a pipeline, for example in Python with OCR, Ollama and a model such as Llama, and compare it with cloud models such as Claude, GPT or Gemini.
- Develop a method that lets the system flag uncertain fields and route them to a person for review.
- Evaluate accuracy per field, processing time and cost per document.
- Formulate recommendations on which approach suits which document volume and which data protection requirements.
What you bring
- Studies in computer science, business informatics, data science or a related field
- Basic Python skills and an appetite for experimenting with language and vision models
- Care in handling data and an eye for clean evaluations
- An interest in processes in logistics, retail or administration
- Good German, as the documents and our clients come from German-speaking countries
Nice to have, not a must: experience with OCR tools, PostgreSQL or workflow tools such as n8n.
What we offer
- Work close to real client projects, not made-up examples
- Subject-matter support from a dedicated contact at vona and regular check-ins; academic supervision is handled by your university
- Access to the AI tools you need for your experiments, such as Claude, ChatGPT, Gemini, Ollama and n8n
- Hours you plan around your semester, working hybrid – in Leipzig and from home
- If it goes well, a chance to keep working together, for example on follow-up projects – a possibility, not a firm promise
How to apply
Use the form at the bottom of this page to tell us a little about yourself and what interests you about document AI. A link to your CV, GitHub or LinkedIn is enough; there’s no need for a cover letter full of stock phrases. You’ll usually hear back from us within a week.