Bachelor's or master's thesis: local language models (m/f/d)
Use your thesis to compare local language models such as Llama via Ollama with cloud models – on quality, cost and data protection for mid-sized firms.
What it’s about
“Are we even allowed to send our data to the cloud?” We hear this question in almost every first conversation with mid-sized companies. The honest answer is often: it depends. Local models keep data in-house but need hardware and maintenance. Cloud models are powerful and quick to set up, but raise questions about data processing agreements and where data flows. In this thesis, you turn “it depends” into a decision guide that people can actually follow.
vona is a web agency from Leipzig that works remote-first – “from Leipzig, for clients everywhere”. Our clients are small and mid-sized companies, established businesses and trades, startups and well-known brands across the German-speaking countries. We use AI as a tool and agree beforehand how success will be measured.
We know both routes from our projects. In the automated document processing we built for a logistics provider, the model runs locally with Ollama and Llama 3 – no document leaves the company, and 85% of documents are still processed fully automatically. Our AI customer advisor for online retail, on the other hand, uses GPT-4o in the cloud and costs around €180 a month to run. One possible research question: at what data volume, level of sensitivity and type of task is a local model the better choice over a cloud model – and how big is the quality gap really? Your own ideas, for instance on hybrid setups or European providers such as Mistral, are welcome.
This position is 100% remote. You can work from anywhere in Germany, Austria or Switzerland.
Your tasks
- Pick two or three typical tasks from mid-sized companies, such as reading documents, summarizing emails or answering questions about internal files.
- Set up local models with Ollama or LM Studio and compare them with cloud models such as Claude, GPT, Gemini or Mistral.
- Design a scoring scheme for answer quality and run the measurements reproducibly, for example in Python.
- Compare costs: hardware and operation on one side, per-request fees on the other.
- Work through the data protection questions and bring everything together in a practical decision guide.
What you bring
- You study computer science, business informatics, data science or something related
- Confidence with Python and the command line
- Curiosity about open-weight models and the drive to get them running yourself
- An interest in GDPR questions – no law degree required
- Good German, as our clients are in German-speaking countries
Nice to have, not a must: experience with Hugging Face, Docker or model quantization.
What we offer
- A topic that comes straight from questions in our client projects
- A dedicated contact who supports you on the subject matter within the company, with regular video check-ins; academic supervision stays with your university
- Access to the tools for your comparisons, including Claude, ChatGPT, Mistral, Ollama and OpenRouter
- Working hours that follow your semester schedule, entirely from home
- If things go well, the possibility of carrying on with us after your thesis, for example on follow-up projects – a possibility, not a promise
How to apply
Fill in the form at the bottom of this page and tell us in a few sentences who you are and why local language models appeal to you. A link to GitHub, LinkedIn or your CV is all we need. You’ll usually hear from us within a week. If you can’t check off everything yet, we’d still like to hear from you.