Browse by type
<img src="https://pydantic.dev/docs/ai/img/pydantic-ai-light.svg" alt="Pydantic AI">
Agents, realtime voice, image generation, embeddings. Every model, every interface, typed end to end.
Pydantic AI is the Python AI SDK: a typed, extensible agent loop with every model a string swap away. The same agent runs everywhere you need it: behind a web frontend, in the terminal, on a voice call, on a durable background queue, in GitHub Actions, or as a plain object you call run() on. Image generation and embeddings come in the same box; Pydantic Graph and Pydantic Evals are separate packages, for typed control flow and for testing agent behavior the way pytest tests code.
Pydantic AI Harness has everything an agent needs for complex, long-running work, snapped on as capabilities, from memory, guardrails, and sub-agents to planning, context management, and storage, up to a complete coding agent.
Pydantic Logfire is the AI observability platform that sees your whole app, not just the LLM calls, and the Pydantic AI Gateway is one key for every model with real-time cost monitoring and budget control; the Gateway self-hosts if you would rather, and our instrumentation is plain OpenTelemetry, so any backend you already run works. Underneath both, genai-prices keeps model pricing current, and Monty is the sandboxed Python interpreter that runs model-written code.
View the complete documentation at pydantic.dev/docs/ai.
From simple typed data extraction to complex, long-running multi-agent collaboration, Pydantic AI and Pydantic AI Harness have got you covered.
A complete coding agent in your terminal: workspace-rooted file access, allowlisted shell, repo orientation, planning, and context management that survives long sessions. Here with web search and a second-opinion advisor snapped on alongside:
uv add pydantic-ai pydantic-ai-harness
from pydantic_ai import Agent
from pydantic_ai.capabilities import WebSearch
from pydantic_ai_harness import Advisor, Coder
agent = Agent(
'anthropic:claude-fable-5',
capabilities=[
Coder(), # files, shell, repo context, sub-agents, context management
WebSearch(), # look up docs and error messages on the web
Advisor('openai:gpt-5.6-sol'), # a second opinion from another model when stuck
],
)
agent.to_cli_sync()
Coder is a regular combined capability, not a black box: use it whole, or use the blocks it bundles directly; the two are equivalent:
capabilities = [
FileSystem('.'), Shell(cwd='.'), RepoContext(), SubAgents(...),
ClearToolResults(), WarnNearLimits(), ToolOutputLimits(), RepairToolArguments(),
]
Run the file and you're chatting with the agent in your terminal. To try it before writing any code, run the exported coder_agent with clai (the Pydantic AI CLI), via uvx:
uvx --with pydantic-ai-harness clai -a pydantic_ai_harness.coder:coder_agent -m anthropic:claude-fable-5
Build this → Coder, from the Harness
Run it on GitHub → GitHub Agentic Workflows, on issues, pull requests or a schedule
Give the agent an output type and tools, and every run comes back validated and typed:
uv add pydantic-ai
from typing import Literal
from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext
class Sentiment(BaseModel):
label: Literal['positive', 'negative', 'neutral']
score: float = Field(ge=-1, le=1)
agent = Agent('openai:gpt-5.6-sol', output_type=Sentiment)
@agent.tool
def recent_reviews(ctx: RunContext[None], product: str) -> list[str]:
"""Fetch recent review snippets for a product."""
return ['The new release fixed everything I complained about!']
result = agent.run_sync('How are people feeling about the Extract app?')
print(result.output)
#> label='positive' score=0.9
The @agent.tool function receives a RunContext that carries your dependencies in; the rest of its signature and its docstring become the tool schema, arguments are validated before your code runs, and the run is guaranteed to return a Sentiment, so your IDE, type checker, and the LLM all agree on the returned type.
Build this → Agents, Function Tools, and Structured Output
Attach TemporalDurability and the same agent runs inside a Temporal workflow under durable execution: every model and tool call becomes a durable activity, so a run working through a background queue survives restarts, failures, and long waits:
uv add "pydantic-ai[temporal]"
from temporalio import workflow
from pydantic_ai import Agent
from pydantic_ai.capabilities import WebFetch, WebSearch
from pydantic_ai.durable_exec.temporal import PydanticAIWorkflow, TemporalDurability
agent = Agent(
'openai:gpt-5.6-sol',
instructions='Research the topic and write a structured brief.',
name='researcher',
capabilities=[WebSearch(), WebFetch(), TemporalDurability()],
)
@workflow.defn
class ResearchWorkflow(PydanticAIWorkflow):
__pydantic_ai_agents__ = [agent]
@workflow.run
async def run(self, topic: str) -> str:
result = await agent.run(f'Write a brief on: {topic}')
return result.output
DBOS and Prefect attach the same way, first-party and co-maintained, with Restate, AWS Lambda, Kitaru, and Airflow integrations besides.
Build this → Durable Execution
Put the same agent on a live voice session, tools and capabilities included:
uv add "pydantic-ai[openai-realtime]"
import asyncio
from pydantic_ai import Agent
from pydantic_ai.capabilities import MCP
agent = Agent(
instructions='You are a helpful voice assistant.',
capabilities=[MCP('https://internal.example.com/mcp')], # capabilities work in voice too
)
@agent.tool_plain
def order_status(order_id: str) -> str:
"""Look up the status of an order."""
return f'Order {order_id}: shipped, arriving Thursday.'
async with agent.realtime('openai:gpt-realtime-2.1').session() as session:
microphone = asyncio.create_task(session.send_audio(microphone_chunks())) # your microphone → the model
speaker = asyncio.create_task(play_audio(session.stream_audio())) # model audio → your speaker
async for part in session.stream_transcripts():
print(f'{part.speaker}: {part.transcript}')
The model calls your tools mid-conversation while it keeps talking, and every session is instrumented; voice is just another frontend, on OpenAI Realtime, Gemini Live, Azure, and xAI Grok Voice.
Build this → Realtime Voice
Generate an image with a dedicated image model, no agent run required:
uv add pydantic-ai
from pathlib import Path
from pydantic_ai import ImageGenerator
generator = ImageGenerator('openai:gpt-image-2')
result = generator.generate_sync('A minimalist logo for a coffee shop called Extract.')
Path('logo.png').write_bytes(result.image.data)
That standalone image API is for when your application decides; when an agent run decides, there is provider-native generation with output_type=BinaryImage for a typed image output, and the ImageGeneration capability with its fallbacks for models that generate no images of their own.
Build this → Image Generation
Any model, one Python API. Virtually every model and provider (OpenAI, Anthropic, Google, Bedrock, Azure AI Foundry, Groq, Mistral, xAI, Ollama, and dozens more), swappable with a string, or through the Pydantic AI Gateway: one key for all of them, with failover and cost monitoring built in. No flagship feature is locked to one vendor.
Typed end to end. Structured outputs, typed dependency injection, typed tools: your IDE, type checker, and coding agent all know what your agent returns, moving whole classes of errors from runtime to write-time. When plain control flow isn't enough, Pydantic Graph brings the same typing to graph-based workflows.
Measured, not vibes. OpenTelemetry-native instrumentation works with any OTel backend; one line lights up Pydantic Logfire for real-time debugging, tracing, and cost tracking backed by genai-prices. Pydantic Evals tests agent behavior the way pytest tests code.
Batteries, composably. One primitive, the [capability](https://py
browse all types & interfaces →
$ claude mcp add pydantic-ai \
-- python -m otcore.mcp_server <graph>