Ollama in Python: Local LLMs for Structured Outputs, Tool Calling, and RAG (2026)
A hands-on 2026 guide to Ollama 0.34 in Python: install the runtime, wire up the ollama client, force JSON-schema outputs with Pydantic, call tools, generate embeddings, and stand up a fully local RAG pipeline against pgvector, with llama.cpp, LM Studio, and vLLM compared.
Ollama is an open-source runtime that lets you download and run large language models locally (Llama 4, Gemma 4, Qwen 3, Mistral, and dozens more) behind a simple HTTP API on localhost:11434. In Python, you install the official ollama package (0.6.2, April 2026), talk to a running Ollama daemon, and get chat completions, JSON-schema structured outputs, tool calls, streaming, and embeddings with no API key, no rate limits, and no data leaving your machine. This guide walks through Ollama 0.34.1 (September 2026) end-to-end for data science and ML engineering workflows.
Honestly, I've been shipping local-LLM features on and off for the past year, and Ollama is the one piece I stopped questioning. I hit this exact stack a few months back on a client project (an on-prem RAG system for a hospital that couldn't send a single row to any cloud), and it just worked. So, let's dig into how to actually use it well.
Ollama 0.34.1 (September 14, 2026) ships JSON-schema structured outputs, tool calling with the tool role, a new web search API, and faster MLX inference on Apple Silicon.
The official ollama Python client exposes chat, generate, embed, and an AsyncClient. It's a thin wrapper over the local REST API on port 11434.
Constrain outputs to a Pydantic model by passing format=Model.model_json_schema() to chat. Ollama compiles that schema into a llama.cpp grammar and forces the tokens to conform.
For local RAG, use ollama.embed() against nomic-embed-text (768 dims, 8k context) or mxbai-embed-large (1024 dims), and store vectors in pgvector 0.8.6+ with HNSW.
Ollama is drop-in OpenAI-SDK compatible via base_url="http://localhost:11434/v1", so you can migrate existing code with a one-line change and no other refactor.
For high-throughput single-model serving on a datacenter GPU, vLLM still wins. Ollama is the right pick for laptop-first development, private data, and multi-model dev loops.
What is Ollama and how does it work?
Ollama is a lightweight local server plus CLI that wraps llama.cpp (and, on Apple Silicon since v0.19, an MLX backend) behind a stable HTTP/JSON API on port 11434. You pull GGUF-quantized weights with ollama pull llama3.3:8b, and the daemon manages loading, memory-mapping, and unloading models as requests arrive. The v0.34 line, released September 14, 2026, adds ChatGPT Desktop integration on macOS, faster structured output on Apple Silicon, a redesigned scheduler that cuts multi-GPU OOM crashes, and a new /api/web_search endpoint backed by Ollama's cloud (free tier included).
The mental model is deliberately simple. One long-running process, one endpoint per capability (/api/chat, /api/generate, /api/embed, /api/tags, /api/ps), and one Modelfile format that describes how a weight file, template, system prompt, and parameters get packaged into a named model. Because Ollama is OpenAI-SDK compatible on /v1, existing tooling (LangChain, LlamaIndex, Instructor, DSPy, the OpenAI Python SDK) talks to it without any Ollama-specific code paths. That is, honestly, the single biggest reason Ollama has become the default local runtime for data scientists in 2026.
Installing Ollama and the Python client
On macOS or Linux, run curl -fsSL https://ollama.com/install.sh | sh. On Windows, use the installer from the release page. That gives you the ollama daemon (auto-started as a background service) and the ollama CLI. Confirm the version, pull a small chat model plus an embedding model, then create a Python virtual environment.
# Runtime
curl -fsSL https://ollama.com/install.sh | sh
ollama --version # -> ollama version is 0.34.1
ollama pull llama3.3:8b # ~4.7 GB, chat model
ollama pull nomic-embed-text # ~274 MB, 768-dim embeddings
# Python side (uv is fastest, pip works fine too)
uv venv .venv && source .venv/bin/activate
uv pip install "ollama>=0.6.2" "pydantic>=2.10" "httpx>=0.28"
The Python package version to install in September 2026 is 0.6.2. It supports Python 3.9+, exposes both a synchronous Client and an AsyncClient, and re-exports module-level helpers (ollama.chat, ollama.generate, ollama.embed) that use a default client pointed at http://127.0.0.1:11434. Point it elsewhere with OLLAMA_HOST=http://gpu-box.internal:11434 or by passing host= to a Client. If you're managing Python environments and reproducible ML projects, our uv package manager guide for data science workflows covers pinned lockfiles and PyTorch backends that pair well with local LLM development.
Chat, streaming, and the async client
The primary entry point is chat(model, messages, stream=False, format=None, tools=None, options=None). Messages follow the OpenAI shape (a list of {"role": "system"|"user"|"assistant"|"tool", "content": "..."} dicts), and the response is a Pydantic-like object where .message.content is the reply text. Turning on stream=True returns an iterator of token chunks so you can render output as it arrives, which is essential for interactive notebooks and any UI where the user is watching.
import ollama
resp = ollama.chat(
model="llama3.3:8b",
messages=[
{"role": "system", "content": "You answer in one short sentence."},
{"role": "user", "content": "Why do we use copy-on-write in pandas 3?"},
],
options={"temperature": 0.2, "num_ctx": 8192},
)
print(resp.message.content)
# Streaming variant
for chunk in ollama.chat(
model="llama3.3:8b",
messages=[{"role": "user", "content": "Explain KL divergence like I know entropy."}],
stream=True,
):
print(chunk.message.content, end="", flush=True)
For anything that fans out (batch classification of a DataFrame column, parallel enrichment of retrieval hits, LLM-as-judge evaluation), use the async client. It uses httpx.AsyncClient under the hood, respects OLLAMA_NUM_PARALLEL on the server side, and lets you saturate a single GPU without spinning up worker processes.
import asyncio, ollama
async def classify(text: str) -> str:
client = ollama.AsyncClient()
r = await client.chat(
model="llama3.3:8b",
messages=[
{"role": "system", "content": "Reply with exactly one label: positive, negative, neutral."},
{"role": "user", "content": text},
],
options={"temperature": 0.0, "num_predict": 3},
)
return r.message.content.strip().lower()
async def main(rows):
return await asyncio.gather(*(classify(t) for t in rows))
labels = asyncio.run(main(["Loved it", "Meh", "Waste of time"]))
print(labels) # ['positive', 'neutral', 'negative']
How does Ollama structured output work?
Since v0.5, Ollama accepts an explicit JSON schema via the format field. Rather than the coarse "json" mode that only guaranteed valid JSON of any shape, the runtime compiles your schema into a llama.cpp GBNF grammar and constrains sampling so every emitted token is legal under that grammar. In practice, this means the model cannot produce a missing required field, a wrong-typed value, or extra keys. Even bad open-weight models cannot break the schema. Compared to prompt-only "please answer in JSON" patterns, this eliminates the entire class of post-hoc parsing failures that used to eat 5–20% of a batch.
The pattern below turns any Pydantic model into an Ollama constraint, then parses the result back into a validated Python object. This is exactly how the Instructor library integrates with Ollama. You can also do it by hand in about ten lines.
from pydantic import BaseModel, Field
from typing import Literal
import ollama, json
class SupportTicket(BaseModel):
priority: Literal["low", "medium", "high", "urgent"]
category: Literal["billing", "auth", "bug", "feature", "other"]
summary: str = Field(max_length=140)
contains_pii: bool
resp = ollama.chat(
model="llama3.3:8b",
messages=[
{"role": "system",
"content": "Extract a support ticket from the message. Return only JSON."},
{"role": "user",
"content": "Hi, my card ending 4242 was charged twice on Monday. Please refund. — Sam"},
],
format=SupportTicket.model_json_schema(),
options={"temperature": 0.0},
)
ticket = SupportTicket.model_validate_json(resp.message.content)
print(ticket.priority, ticket.category, ticket.contains_pii)
A few real-world lessons the docs don't spell out. First, keep temperature at 0.0 for extraction; grammar-constrained sampling still benefits from a peaked distribution when you want repeatable behavior. Second, the model can still be semantically wrong (the schema guarantees shape, not correctness), so validate downstream with Pandera checks or business rules. And third, structured outputs cost tokens; complex nested schemas expand the grammar and slow generation by 10–30%. For deeper coverage of the design tradeoffs across libraries, our comparison of Instructor vs Outlines vs Pydantic AI for structured LLM outputs maps out when to reach for each.
Tool calling with the tool role
Ollama's chat endpoint accepts a tools= argument shaped like the OpenAI function-calling schema. When the model decides a tool is needed, it returns message.tool_calls (a list of {name, arguments} pairs), and you execute the function, append the result as a {"role": "tool", "content": ...} message, and call chat again to let the model incorporate the result. The tool role, added in the 2026 client updates, is now the canonical way to feed observations back into the loop.
import ollama, json
from datetime import datetime, timezone
def get_now(tz: str = "UTC") -> str:
return datetime.now(timezone.utc).isoformat()
TOOLS = [{
"type": "function",
"function": {
"name": "get_now",
"description": "Return the current UTC timestamp.",
"parameters": {"type": "object", "properties": {"tz": {"type": "string"}}},
},
}]
messages = [{"role": "user", "content": "What time is it right now?"}]
resp = ollama.chat(model="llama3.3:8b", messages=messages, tools=TOOLS)
for call in resp.message.tool_calls or []:
args = call.function.arguments
result = get_now(**args)
messages.append(resp.message)
messages.append({"role": "tool", "name": call.function.name, "content": result})
final = ollama.chat(model="llama3.3:8b", messages=messages, tools=TOOLS)
print(final.message.content)
Reliability of tool calling depends heavily on the underlying weights. In 2026, Llama 3.3 8B, Qwen 3.5, and Gemma 4 26B are all reliable enough that agent loops which used to require Claude Sonnet now work locally. Almost every Ollama release in 2026 (v0.17.5, v0.18.3, v0.20.1, v0.20.7, v0.22.1) shipped tool-call fixes, so if you're stuck on stale behavior, upgrade the daemon before you rewrite the prompt. For stateful multi-step agents with checkpointing and human review, see our guide to LangGraph for stateful LLM agents, which plugs into Ollama just as cleanly as it does OpenAI.
Embeddings and a local RAG pipeline with pgvector
ollama.embed(model, input) accepts a string or a list of strings and returns a list of float vectors. The two most popular embedding models in 2026 are nomic-embed-text (768 dims, 8192-token context, ~274 MB) and mxbai-embed-large (1024 dims, ~669 MB). Qwen3 Embedding is the go-to for Chinese-heavy corpora. A single RTX 4090 with nomic-embed-text and OLLAMA_NUM_PARALLEL=4 exceeds 1,000 chunks per second batched, and that's enough that you rarely need a cloud embedding vendor for corpus sizes under about 10 million chunks.
Below is a minimal, fully-local RAG pipeline. Embed a corpus, store in pgvector 0.8.6+ with an HNSW index, retrieve for a query, then answer with a Llama chat model. It's deliberately dependency-light so you can see every moving part.
import ollama, psycopg
from pgvector.psycopg import register_vector
DOCS = [
"Copy-on-write in pandas 3 makes chained assignment safe.",
"Polars uses Apache Arrow columnar layout for zero-copy IO.",
"DuckDB executes SQL directly against Parquet without loading.",
]
conn = psycopg.connect("dbname=rag user=postgres", autocommit=True)
register_vector(conn)
conn.execute("CREATE EXTENSION IF NOT EXISTS vector")
conn.execute("""
CREATE TABLE IF NOT EXISTS chunks (
id serial PRIMARY KEY,
content text NOT NULL,
embedding vector(768)
)""")
conn.execute("""
CREATE INDEX IF NOT EXISTS chunks_hnsw
ON chunks USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64)""")
# Embed corpus in one batched call
vecs = ollama.embed(model="nomic-embed-text", input=DOCS).embeddings
with conn.cursor() as cur:
cur.executemany(
"INSERT INTO chunks (content, embedding) VALUES (%s, %s)",
list(zip(DOCS, vecs)),
)
# Query
q = "What speeds up analytical queries on Parquet files?"
qvec = ollama.embed(model="nomic-embed-text", input=q).embeddings[0]
rows = conn.execute(
"SELECT content FROM chunks ORDER BY embedding <=> %s LIMIT 3",
(qvec,),
).fetchall()
context = "\n".join(r[0] for r in rows)
answer = ollama.chat(
model="llama3.3:8b",
messages=[
{"role": "system", "content": "Answer only from CONTEXT. Cite by index."},
{"role": "user", "content": f"CONTEXT:\n{context}\n\nQUESTION: {q}"},
],
options={"temperature": 0.1},
)
print(answer.message.content)
Two production-critical details. Use pgvector 0.8.2 or newer; earlier versions have an unpatched buffer overflow in parallel HNSW builds (CVE-2026-3172). And always use the plural /api/embed endpoint or the ollama.embed() helper. The deprecated singular route is single-input only and much slower. If you want to compare pgvector against dedicated stores like LanceDB, Qdrant, or Milvus for your workload, our Python vector databases comparison covers latency, cost, and hybrid search tradeoffs. Once you have retrieval working, layer a cross-encoder on top. Our guide to reranking for RAG with BGE, Jina, and ColBERT shows how to lift retrieval quality by 20–40 nDCG points with a second-stage ranker.
Ollama vs llama.cpp vs LM Studio vs vLLM
The four options overlap but optimize for different jobs. Ollama is developer-first: one binary, a package manager for weights, an OpenAI-shaped API, and no GUI in the way. llama.cpp is the underlying inference engine that Ollama drives; you pick it if you want direct control over the C++ layer or need to ship a single static binary. LM Studio is the GUI-first option, useful for non-developers evaluating models. vLLM is a datacenter serving engine, with much higher throughput on a single big GPU but heavier to deploy and less friendly for local dev loops.
Dimension
Ollama
llama.cpp
LM Studio
vLLM
Primary interface
CLI + REST API
C++ / CLI
Desktop GUI
Python + REST
Best for
Local dev loops, private data
Embedded / static binary
GUI evaluation
Multi-user GPU serving
OpenAI API compat
Yes (/v1)
Yes (server mode)
Yes
Yes (native)
Structured outputs
JSON schema (v0.5+)
GBNF grammar
JSON schema
Outlines / xgrammar
Tool calling
Yes, tool role
Model-dependent
Yes
Yes
Throughput (single 4090)
~60 tok/s Llama-8B
~65 tok/s Llama-8B
~55 tok/s
~180 tok/s batched
Concurrency model
OLLAMA_NUM_PARALLEL
Manual
Limited
Continuous batching
Quantization
GGUF Q4/Q5/Q8, MLX
GGUF, all quant levels
GGUF
AWQ, GPTQ, FP8
License
MIT
MIT
Proprietary (free tier)
Apache 2.0
Practical rule of thumb: use Ollama for laptops, private data, notebooks, and iteration on prompt/model design. Move to vLLM once a specific model needs to serve many concurrent users on a datacenter GPU with strict latency budgets. If you find yourself operating a fleet of models rather than a single one (different customers, different fine-tunes), a gateway like LiteLLM in front of both Ollama and cloud providers keeps client code identical. Our LiteLLM gateway guide covers cost tracking and routing.
Which models should you actually run?
The right model depends on your VRAM budget and task. As of September 2026, five profiles cover about 90% of Python data science needs.
Chat and general reasoning
llama3.3:8b (Q4_K_M, ~4.7 GB) is the sweet spot for laptops with 16 GB unified memory or an 8 GB discrete GPU. Bump to llama3.3:70b (Q4_K_M, ~40 GB) when you have a 48 GB card or two consumer GPUs. For agentic tool use, qwen3.5:14b and gemma4:26b both fixed the tool-call regressions that plagued mid-2026 releases.
Code generation
qwen3-coder:30b on Ollama's cloud (or local, if you have the VRAM) is currently the highest-scoring open code model on HumanEval+ and LiveCodeBench, and Ollama's fall 2026 update made its tool calling faster and more reliable.
Embeddings
nomic-embed-text (768 dims, 8k tokens) is the default. Switch to mxbai-embed-large (1024 dims) for slightly higher retrieval recall at the cost of about 30% more storage. Both are Apache-licensed.
Vision
Ollama supports Alibaba's qwen3-vl for OCR, chart understanding, and image-grounded Q&A. Pair it with Docling for PDF and DOCX parsing when your RAG corpus is heavily documents.
Production tips: GPU, concurrency, and observability
Ollama is optimized for local dev, but you can absolutely serve it in production if you know the levers. Set OLLAMA_NUM_PARALLEL to enable request interleaving on one loaded model (2–4 is right for consumer GPUs, 8+ for datacenter cards). OLLAMA_MAX_LOADED_MODELS controls how many distinct models the daemon keeps resident; increase it to avoid unloading/reloading between requests, but watch VRAM.
The v0.34 scheduler rewrite dramatically cuts multi-GPU OOM crashes, so if you're running two or more cards, upgrade before you invest in workaround scripts. Turn on OLLAMA_FLASH_ATTENTION=1 to get FlashAttention on any GPU with compute capability 6.x or higher (v0.31.2 broadened support). For observability, the daemon exports Prometheus metrics on /metrics. Pair them with a tracing layer around your Python client. Our comparison of Langfuse vs LangSmith vs Arize Phoenix covers what to trace and how to catch latency regressions before users do.
Finally, put a real HTTP client in front of the daemon in production. The default requests path used by some third-party wrappers has no connection pooling. The official ollama package uses httpx with keep-alive, which is what you want. Set client-side timeouts, retry only on 5xx and timeouts (not 4xx), and log the model tag with every span so a rollback is a one-line grep. Consult the official ollama-python repository and the Ollama blog for the current release notes and model catalog before pinning a version.
Frequently Asked Questions
Is Ollama free to use?
Yes. Ollama itself is MIT-licensed and free to run locally with no account. The optional cloud service (used for the ollama.com-hosted models and the web-search API) has a generous free tier and paid tiers for higher rate limits, but nothing about the local runtime requires an account or a network connection.
Does Ollama work offline?
Yes, once you've pulled the model weights. After ollama pull completes, the daemon runs inference entirely from local disk and can operate on an air-gapped machine. Only ollama pull, ollama push, and the new web-search API require network access.
Can Ollama replace the OpenAI API?
For most feature parity, yes. Ollama exposes an OpenAI-compatible endpoint at http://localhost:11434/v1, so you can point the official OpenAI Python SDK at it with a one-line base-URL change. The gap is quality on frontier tasks (GPT-5 or Claude Sonnet still beat open weights on the hardest reasoning benchmarks), but for extraction, classification, RAG, and most agent loops, a well-chosen Ollama model is enough.
How much RAM or VRAM do I need to run Ollama?
Rule of thumb: an 8B model at Q4 fits in about 5 GB of VRAM (or unified memory on Apple Silicon), a 14B in about 9 GB, a 26B in about 16 GB, and a 70B at Q4 in about 40 GB. Add roughly 2 GB overhead for KV cache at 8k context. If the model doesn't fit in VRAM, Ollama transparently offloads layers to CPU, which drops throughput 3–10x. Plan hardware accordingly.
How do I run Ollama on a GPU?
On Linux with an NVIDIA GPU, install the NVIDIA driver and container toolkit; Ollama will detect CUDA automatically. On macOS, the MLX backend (since v0.19) uses the Apple Silicon GPU by default with no configuration. On Windows, install the latest NVIDIA driver and Ollama will use CUDA. Confirm with ollama ps. The PROCESSOR column shows gpu or cpu per running model.
A FastAPI developer's tour of Modal in 2026: decorators, per-second GPU pricing, Volumes, @app.cls with @modal.enter, FastAPI endpoints, spawn_map for batch fan-out, and the cold-start tuning knobs that decide whether Modal saves money or bleeds it.
Zarr-Python 3 brings full v3 spec support, an async core, and chunk sharding for cloud object stores. A data-engineering walkthrough with chunking rules, migration steps, and pipeline tests you can actually run.
Benchmark Cohere Rerank 3.5, BGE v2-m3, Jina Reranker v2, and ColBERT v2 for RAG in Python. Runnable code, NDCG@10 results, latency, and $/1M queries so you can pick the right reranker.