Skip to content

Type a term, e.g. API, Shopware or SEO.

    ← All posts
    Artificial Intelligence

    By René Mattis · vona ·

    Reducing Ongoing AI Costs: Why Not Every Step Needs AI

    Why AI automations often cost more than necessary, when conventional code is the better choice, and how an orchestrated workflow cuts costs significantly – using document management as an example.

    An automation is running, the team is happy, and yet the bill grows every month: from the AI provider, from n8n or Make, sometimes both. We see this regularly in practice. The reason is rarely the model itself and almost always the planning: AI does jobs that ordinary program code would do faster, cheaper and more reliably. This article shows where ongoing AI costs come from, how to decide where AI is really needed, and works through a typical example: processing incoming documents.

    Where ongoing costs come from

    With AI processes you pay in two places:

    • The AI provider, by usage, measured in tokens, i.e. pieces of text. You pay for everything the model reads and everything it writes. Output tokens are usually five times as expensive as input tokens. Images cost tokens too, and a scanned page as an image is considerably more expensive than the same content as text.
    • The automation platform, and here providers differ widely: n8n Cloud bills each execution of a complete workflow, no matter how many steps it has. Since August 2025, Make bills in credits, basically one per module run, with AI modules using more. Zapier counts every successful action step as a task.

    This leads to three typical cost drivers: too many calls (each sub-step sent to the AI separately), too much context (the whole document sent again with every call) and models that are too large (the default model for everything, including jobs a small model – or no model at all – could handle).

    The real problem: AI as an all-purpose tool

    Many automations are created like this today: someone describes the desired workflow to an AI assistant, and the assistant designs it. The result works, but it is often built the way a language model thinks: every step is handed to a language model. Read the date from an invoice? AI. Check the total? AI. Check whether an IBAN is valid? AI. Name the file? AI.

    The promise of “100% automated” works similarly. That can be true, but it does not mean that 100% of the work should be done by AI. Automation simply means that no person has to step in. Whether a step is handled by a rule, by program code or by AI is a separate decision – and without proper planning, i.e. without orchestration, that decision is often never made.

    Yet for such jobs AI has real drawbacks: every call costs money, takes seconds instead of milliseconds and does not always return the same result. With calculations it can even be wrong (hallucination). Conventional code, on the other hand, costs practically nothing per run once it has been written, is fast and always behaves the same way.

    Rule of thumb: code or AI?

    Task Better with code Better with AI
    Data in a fixed format (date, amount, IBAN, invoice number) ✓ patterns, check digits, format rules
    Calculating and checking (totals, tax, matching against orders) ✓ always exact
    Known senders and consistent layouts ✓ templates and rules
    Structured data (XML, JSON, interfaces) ✓ read directly
    Understanding free text, recognising intent ✓
    Unknown layouts, handwritten notes, inconsistent wording ✓
    Summarising, writing, translating ✓
    Classifying unclear cases ✓ with a human check

    So the basic rule is: code for everything that can be described as a rule, AI for everything that requires understanding language. And when AI, then the smallest model that solves the task reliably. The providers point the same way: Anthropic describes starting with the small model and upgrading only if needed as one approach in its documentation, and OpenAI’s cost optimisation guide explicitly recommends making fewer requests, using fewer tokens and choosing a smaller model.

    Example: processing documents automatically

    Take a typical case: a company receives invoices, delivery notes and contracts by email and post. They need to be recognised, read, checked, filed and summarised where necessary. Here is what an orchestrated workflow looks like, step by step:

    1. Receive and detect the format (code): What kind of file arrives – PDF, image or XML? A simple check, no AI needed.
    2. Read e-invoices directly (code): Since 1 January 2025, businesses in Germany must be able to receive e-invoices, and their share is growing with transition periods running until the end of 2027. XRechnung and ZUGFeRD contain the invoice data as structured XML; for ZUGFeRD, according to the German Federal Ministry of Finance, this XML is authoritative, not the image of the PDF. Code reads this data directly, for example with the Python library factur-x. No text recognition and no AI required.
    3. Get the text (code): Digital PDFs already contain their text; pdfplumber extracts it, including tables. Scanned documents get a text layer with OCRmyPDF (based on the Tesseract OCR engine), running locally with no per-page cost. For difficult originals there are specialised services such as Mistral OCR 4.1 (list price USD 4 per 1,000 pages) or tools like Docling that convert documents, including their structure, into structured text. The key point: from here on, you work with text, not images.
    4. Determine the document type (code first, then AI): Sender address, keywords such as “invoice” or “delivery note” and known layouts classify most documents by rule. Only what is left goes to a small language model – and only the beginning of the text, not the whole document.
    5. Extract the data (code first, then AI): For known senders, templates with patterns are enough, for example with invoice2data. dateparser recognises dates in various notations, RapidFuzz finds fuzzy matches such as slightly different company names. Only for unknown layouts does a small model extract the fields, with a fixed output format so that code can process the result directly.
    6. Check (code): Do net amount and tax add up to the total? Does the order number match an open order? Is the IBAN valid? These are calculations, and code calculates exactly. Anything that fails the check goes to a person (human in the loop).
    7. File and hand over (code): File name, folder, transfer to ERP or accounting – all by fixed rules.
    8. Summarise (AI, targeted): A contract or a long letter needs to be explained in three sentences? That is a genuine language task, and AI is strong here. But only where a summary is needed, not for every invoice.

    The result: AI is used at just two or three points, and there with short text excerpts instead of whole page images.

    A note on libraries: if you want to process PDFs with PyMuPDF, check the licence. PyMuPDF is licensed under the AGPL; for closed-source software or an online service you need a commercial licence from the vendor or an alternative such as pdfplumber (MIT licence).

    What this means in numbers

    An example with 2,000 documents a month, two pages each. The assumptions are deliberately simple: a page as an image corresponds to about 1,600 tokens, an instruction to the model to about 1,000 tokens, a response to about 300 tokens. We use the providers’ list prices (as of October 2026, in US dollars, excluding tax) and only the pure model costs.

    Variant Setup Model costs per month
    A: everything via AI, mid-size model Each document goes to Claude Sonnet 5.5 (USD 2/10 per 1M tokens) four times as an image: classify, extract, check, summarise about USD 91
    B: everything via AI, small model The same workflow with Claude Haiku 5.5 (USD 0.10/0.50 per 1M tokens) about USD 4.60
    C: orchestrated Local OCR, rules and templates first; about a third of documents need a small model for extraction (as text), one in ten a summary under USD 1
    D: orchestrated, local model Like C, but the small model runs on your own hardware (see below) USD 0 for the model, but hardware, power and maintenance instead

    The calculation shows two things. First: the choice of model alone makes a difference by a factor of 20. Second: the biggest saving in C does not come from the price alone but from the fact that most steps no longer go to a model at all. On top of that come benefits the table does not show: code steps run in milliseconds, always return the same result and calculate exactly. And because platforms bill per execution, module or step depending on the provider, fewer AI steps often also mean lower platform costs – with Make, for instance, Make’s own AI modules use more credits than ordinary modules.

    To be fair: with a few hundred documents a month, pure model costs are manageable even in variant A. Costs rise with volume, with large models as the default and with long documents. And anyone who only looks at model costs overlooks the cost of errors: an incorrectly extracted total that only surfaces in accounting costs more than any token.

    More levers against rising costs

    • The right model for each task: small models for routine work, large ones only for difficult cases. An upstream LLM gateway can distribute this automatically.
    • Caching: providers cache recurring parts of a request, such as long instructions or templates. At Anthropic, reading from the cache costs only a fraction of the normal input price, and the same applies at OpenAI.
    • Batch processing: anything that does not have to be finished immediately, such as nightly processing, costs 50% less via the batch interfaces of Anthropic, OpenAI and Google.
    • Less context: send only the excerpt the model really needs, and text instead of images wherever possible.
    • Fixed output formats: when the model returns a fixed format, follow-up questions and correction loops disappear.
    • Budget limits and per-step measurement: only if you know what each step costs can you optimise in a targeted way. Limits protect against outliers, for example when an error triggers a loop.

    You will find more on costs and an example calculation for ongoing operation on our pricing page.

    Running open models locally

    Another option: the language model does not run at the provider at all but on your own hardware, as a local model. Then there are no per-call costs, and the documents never leave the company. Especially for the steps in our example – classifying documents and extracting fields from text – a large model is often not needed.

    The tools: The easiest is Ollama (MIT licence). It downloads models with a single command and can force responses into a fixed JSON format so that code can process them directly. Underneath it runs llama.cpp (MIT). For a server handling many simultaneous requests there is vLLM (Apache-2.0). LM Studio is an interface for trying things out; since July 2025 it has been free for work use as well, but it has its own terms of use and on the Mac only runs on Apple chips.

    The models: With open weights and a permissive licence, there is now a good choice of models in sizes that run on a normal workstation or a small server, all under Apache-2.0 or MIT and therefore usable commercially:

    • Ministral 3 by Mistral (3, 8 and 14 billion parameters, German explicitly supported, with JSON output),
    • Gemma 4 by Google (including 12 and 31 billion parameters),
    • Qwen3.5 by Alibaba (including 4, 9 and 27 billion parameters),
    • Granite 4.2 by IBM (3, 8 and 30 billion parameters, German supported),
    • Phi-4 by Microsoft (14 billion parameters, MIT),
    • gpt-oss-20b by OpenAI, which according to the vendor runs within 16 GB of memory.

    It is still worth reading the licence: not every “open” model is free to use. Meta’s Llama 4, for example, has its own licence which, for the multimodal models, excludes companies with their principal place of business in the EU.

    The hardware: As a rule of thumb, a model in the usual 4-bit version needs a little over half a gigabyte of memory per billion parameters, plus room for the context. In Ollama, an 8-billion model takes about 6 GB, a 14-billion model about 9 GB, a 24-billion model about 15 GB. On older computers without a suitable GPU or Apple chip this runs noticeably slower, though for a nightly batch it is often still enough.

    Testing whether it is good enough: Only a test with real data shows whether an open model is sufficient for your task. A proven approach: pick 50 to 100 typical documents, record the correct values for each, then compare field by field what the local model and, for comparison, a small cloud model deliver. Measure accuracy, time per document and the cases where the model is unsure. If the local model keeps up, the decision is easy. If not, a slightly larger model or better text preparation often helps before switching to the cloud.

    An honest calculation: A local model is not free either. Hardware, power, setup, updates and monitoring cost money and time. It pays off above all at high volume, with sensitive data or when suitable hardware is already available.

    Data protection and the EU AI Act: For data protection, running locally is a real advantage. The data goes to no AI provider, there is no processing agreement needed for it and no transfer to third countries. In its guidance on AI, the German Data Protection Conference states that “technically closed systems” are preferable from a data protection perspective. For the EU AI Act, by contrast, location changes nothing: obligations depend on your role and the purpose of use, not on where the model runs. If you use a local model, you still have to support AI literacy in your team, meet transparency obligations and, for high-risk uses such as recruitment, fulfil the strict requirements. The regulation’s exception for systems under free and open-source licences (Art. 2(12)) is no free pass either: it does not apply to high-risk systems or to the cases in Art. 5 and 50, and whether it applies to a company’s own business use at all has not been conclusively settled. More on this on our EU AI Act page.

    Conclusion

    AI is a powerful tool, but no substitute for proper planning. The cheapest and most reliable automation is the one in which every step is done with the right means: rules and code for everything that can be described, AI for what needs language understanding, and a person for the cases where something is unclear. That keeps costs predictable – and the results get better.

    A second real-world example with actual figures – AI product images via a custom app instead of a platform subscription – is in the article AI product images: why the subscription cost more than our own solution. Our case study Automated document processing for a logistics company, where the language model runs locally within the company, shows what this looks like in practice. If you already run AI processes or automations and the costs keep rising, let’s look at them together: we analyse existing workflows, show which steps run cheaper and more reliably as code, and implement the optimisation on request. More under AI consulting or directly in a free initial consultation.

    As of October 2026. Prices are the providers’ list prices according to their pricing pages (Anthropic, OpenAI, Google, Mistral, n8n, Make, Zapier), licences and sizes according to the model cards and project pages, retrieved on 10 October 2026; they change regularly. Data Protection Conference quote: guidance “Künstliche Intelligenz und Datenschutz”, version 1.0 of 6 May 2024.

    ← Back to overview

    Newsletter

    News from the vona workshop.

    AI tools we actually use, our weekly AI review & insights from our projects – short, practical and without spam.