LlamaIndex
Open-source framework for Python and TypeScript for building AI apps on your own documents, plus the LlamaParse cloud service for parsing.
What it does
LlamaIndex is a library for Python and TypeScript that connects language models to your own data. Connectors load content from files, databases and other sources, indexes and vector stores make it searchable, and query engines return answers with the right context (RAG). On top of that come agents that call tools and workflows for multi-step processes with branching, retries and human review. The framework works with cloud models such as OpenAI’s as well as with local models. The paid cloud service LlamaParse converts PDFs, spreadsheets and scans into Markdown or JSON, extracts fields according to a schema and sorts documents into categories.
Who it suits
Development teams that want to build a chatbot over manuals, a search across contracts or invoice processing themselves. Without programming skills you won’t get far.
Data protection for businesses
The framework runs on your own servers; where data goes depends on the model provider you choose. For LlamaParse there is an EU region alongside the US region. Uploaded files are cached for 48 hours, and caching can be switched off with a parameter. Inputs are not used for training by default; that only changes if an admin actively opts into a data-sharing programme. You request a data processing agreement via an online form. Deployment in your own cloud and HIPAA agreements are available only in the Enterprise plan.
Limits
LlamaParse bills in US dollars via credits. The number and size of indexes are capped per plan; the free plan allows 5 indexes with 50 files each. Answer quality depends heavily on the connected language model.