AI Engineer: build a real AI agent, module by module
Each module combines clear theory, commented code examples and hands-on exercises. Once you finish them, all the deliverables come together into a production-ready agent system.
- modules
- 10
- deliverables
- 10
- final project
- 1
- estimated duration
- ~10 wks
Final project: Nexus Support Agent
A multi-agent customer support system. End-to-end: RAG, episodic memory, tool calling, human-in-the-loop, model routing, observability and continuous evaluation in CI.
Architecture
| Input | Core Agents | Infrastructure |
|---|---|---|
| Chat API (FastAPI) | Orchestrator | pgvector + RAG |
| Classifier Agent | SupportAgent (ReAct) | Redis (session) |
| MemoryManager | CriticAgent | LangSmith (tracing) |
| Guardrails | EscalationRouter | Eval pipeline (CI) |
Which component each module contributes
- M1 →
LLMClient - M2 →
PromptLoader - M3 →
SupportAgent - M4 →
Orchestrator - M5 →
RAG+Memory - M6 →
EscalationRouter - M7 →
ToolRegistry - M8 →
ModelRouter - M9 →
DebugToolkit - M10 →
Observability
Suggested schedule
One week per module, with progressive integration into the final project.
| Weeks | Phase | What you build |
|---|---|---|
| Weeks 1-2 | Foundations | LLMClient + PromptLoader |
| Weeks 3-5 | Architecture | Agent + Orchestrator + RAG |
| Weeks 6-7 | Production | HITL + Tool Layer |
| Weeks 8-9 | Advanced | Optimization + Debug |
| Week 10 | LLMOps | Observability + Integration |
Module 1 · Phase 1 · LLMs & Prompt Engineering
LLM Fundamentals
Architecture, inference and model selection in production
What is an LLM and how does it generate text?
A Large Language Model is a neural network trained to predict the next token given a context. It doesn't "understand" the way we do: it learned statistical patterns from trillions of tokens of text. Every time it generates a word, it computes a probability distribution over the vocabulary and samples from it.
Imagine you completed millions of "finish this sentence" exercises. With enough practice, you develop an intuition for which words tend to follow which. An LLM does something similar, but at massive scale and with far more complex patterns.
The Transformer architecture: what you need to know
You don't need to implement a Transformer, but you do need to understand its practical implications:
Key concepts and their impact in production
- Self-attention: each token "attends" to every other token in the context. Implication: the model can relate information that is far apart in the text.
- Context window: the maximum number of active tokens. GPT-4: 128K, Claude 3.5: 200K, Gemini 1.5 Pro: 1M. More context = more cost and latency.
- Decoder-only: modern models (GPT, Claude, Gemini) only generate; they don't have a separate encoder. They process the whole context every time they generate a token.
- Tokenization (BPE): text is split into sub-words. "tokenization" can be 3-4 tokens. The real cost of a call depends on the number of tokens, not words or characters.
Inference parameters: the control dial
When you call the API, these parameters determine how the model samples its response:
temperature
0 = deterministic (always the most likely token). 1 = more varied. For agents in production: use 0–0.3. For creative generation: 0.7–1.0.
top_p (nucleus sampling)
Only considers the tokens whose cumulative probability reaches p. top_p=0.9 ignores the least likely 10% of tokens. An alternative to temperature; they aren't used together.
max_tokens
The token limit for the response. It directly affects cost and latency. For structured responses (JSON), a low limit prevents unexpectedly truncated responses.
stop_sequences
The model stops generating when it hits this string. Useful for delimiting outputs: ["</response>", "###"]. More reliable than max_tokens for structured outputs.
Using temperature=0 doesn't guarantee identical outputs. LLMs can vary even at temperature=0 because of differences in hardware and parallelism. For exact reproducibility, store the full input and the output.
Model selection: the most important trade-off
Selection guide for production, 2025
- Claude 3 Haiku / GPT-4o-mini: classification, routing, simple extraction. ~$0.25/M tokens. Latency: <1s. Use it when errors have low impact.
- Claude 3.5 Sonnet / GPT-4o: reasoning, generation, complex tool calling. ~$3/M tokens. The ideal balance for most agents in production.
- Claude 3 Opus / GPT-4-turbo: deep analysis, high-impact decisions. ~$15/M tokens. Only when quality is critical and cost is secondary.
- Mistral / LLaMA 3 (self-hosted): sensitive data, regulatory compliance, cost at extreme scale. Requires your own infrastructure.
70-80% of the queries in a support system are simple. Automatically classifying complexity and using Haiku for simple cases can cut total cost by 60% with no noticeable impact on quality.
LLMClient: the system's base wrapper
The right pattern is not to call the SDK directly from each agent. You create a centralized wrapper that handles automatic retries, structured logging, token counting and model selection.
from anthropic import Anthropic
from tenacity import retry, stop_after_attempt, wait_exponential
from enum import Enum
import structlog, time
log = structlog.get_logger()
class ModelTier(Enum):
FAST = "claude-3-haiku-20240307" # barato y rápido
STANDARD = "claude-3-5-sonnet-20241022" # balance ideal
POWERFUL = "claude-3-opus-20240229" # máxima calidad
class LLMResponse:
text: str
input_tokens: int
output_tokens: int
cost_usd: float
latency_ms: float
class LLMClient:
def __init__(self):
self.client = Anthropic()
self.cost_per_token = {
ModelTier.FAST: (0.00025, 0.00125), # (input, output) por 1K tokens
ModelTier.STANDARD: (0.003, 0.015),
ModelTier.POWERFUL: (0.015, 0.075),
}
@retry(stop=stop_after_attempt(3),
wait=wait_exponential(multiplier=1, min=1, max=10))
def call(
self,
messages: list[dict],
model: ModelTier = ModelTier.STANDARD,
temperature: float = 0.3,
max_tokens: int = 1024,
trace_id: str = None,
) -> LLMResponse:
start = time.time()
response = self.client.messages.create(
model=model.value,
messages=messages,
temperature=temperature,
max_tokens=max_tokens,
)
latency = (time.time() - start) * 1000
cost = self._calculate_cost(model, response.usage)
# Log estructurado para observabilidad
log.info("llm_call",
trace_id=trace_id,
model=model.value,
input_tokens=response.usage.input_tokens,
output_tokens=response.usage.output_tokens,
cost_usd=round(cost, 6),
latency_ms=round(latency, 1),
)
return LLMResponse(
text=response.content[0].text,
input_tokens=response.usage.input_tokens,
output_tokens=response.usage.output_tokens,
cost_usd=cost,
latency_ms=latency,
)
def _calculate_cost(self, model, usage) -> float:
inp, out = self.cost_per_token[model]
return (usage.input_tokens/1000*inp) + (usage.output_tokens/1000*out)
The @retry decorator from tenacity automatically handles rate limits (429) and transient API errors with exponential backoff. Without it, any transient failure breaks the agent's flow.
Resources
anthropic-sdk docs, tenacity, structlog, tiktoken
Temperature benchmark
Understand empirically how temperature affects the output before choosing the value for production.
- Write a prompt that asks to classify the sentiment of a sentence (positive/negative/neutral)
- Run the same call 10 times with temperature=0. Does it always give the same result?
- Repeat with temperature=0.5 and temperature=1.0. What changes?
- Record the cost and latency of each call. Does temperature affect cost?
- Conclusion: which temperature would you choose for the project's intent classifier?
Token counting and cost estimation
Before designing the system, know how much each call will cost.
- Write the support agent's system prompt (initial draft, ~200 words)
- Use
tiktokento count how many tokens it takes - Simulate 1000 five-turn conversations: calculate the total cost with Haiku vs Sonnet
- What percentage of the cost comes from the system prompt vs the history?
- Document which model you would choose for the intent classifier, and why
Module deliverable
Goes into the final project: LLMClient wrapper. A production-ready Python class that wraps every API call. Every agent in the project will use this wrapper, never the SDK directly. The LLMClient is the base layer of the system. In M8 (Model Router) it will be extended to select the model dynamically per query, instead of receiving it as a fixed parameter.
Module 2 · Phase 1 · LLMs & Prompt Engineering
Advanced Prompt Engineering
System prompts, constraints, controlled CoT and versioning
The system prompt is the agent's constitution
The system prompt is not an "initial instruction". It is the document that fully defines who the agent is, what it can do, how it should behave and when it should ask for help. An agent without a well-structured system prompt is an unpredictable agent.
6-section structure (all required)
- IDENTITY: name, purpose, personality. Defines the "who" of the agent.
- CAPABILITIES: an explicit list of what it can and CANNOT do. The "cannots" matter as much as the "cans".
- CONTEXT: dynamic variables from the environment: user, state, available tools.
- BEHAVIOR RULES: how to act in specific situations, with explicit edge cases.
- OUTPUT FORMAT: the exact structure, length and channel of the response.
- ESCALATION: exact criteria for handing off to a human.
Whatever isn't explicit in the system prompt, the model makes up. Every expected behavior has to be specified. Ambiguity in the prompt is the source of 80% of agent bugs.
Controlled Chain-of-Thought: separating reasoning from the answer
CoT (Chain-of-Thought) improves reasoning quality, but in production we don't want to show the internal process to the user. The right pattern is to separate the two:
Uncontrolled CoT
- The user sees the internal reasoning
- It exposes logic that can be manipulated
- It needlessly increases output tokens
- It makes the final answer harder to parse
CoT with a separate <thinking>
- The reasoning stays in internal logs
- It allows debugging without exposing it to the user
- The final output is clean and parseable
- You can monitor the quality of the reasoning
Few-shot with negative examples
Positive examples teach the expected behavior. Negative examples are just as critical: they show the model exactly what to avoid. Without them, the model can fall into "reasonable but wrong" answers.
Including only positive examples in the few-shot. The model learns "what to do" but not "what NOT to do". The most frequent edge cases and failures should appear as explicit negative examples.
Prompts as code: versioning and testing
A prompt that changes without control is a silent regression. The same discipline we apply to code applies to prompts:
Prompt deployment pipeline
- A git PR with the prompt change + the rationale in the description
- Automatic evaluation in CI against the base test set
- Deploy to staging → 10% of traffic → 48h of monitoring
- If metrics are OK → promote to 100%. If they degrade → automatic rollback
Complete system prompt with controlled CoT
IDENTITY:
Eres SupportBot, asistente de atención al cliente de Nexus.
Objetivo: resolver consultas de soporte con empatía y precisión.
Tono: cercano, claro, sin jerga técnica innecesaria.
CAPABILITIES:
✓ Puedes: get_order_status, create_ticket, send_notification, schedule_callback
✗ NO puedes: modificar precios, eliminar cuentas, acceder a datos de pago
CONTEXT:
Usuario: {{user_name}} | Plan: {{plan_name}} | Estado: {{account_status}}
Canal: {{channel}} | Herramientas: {{available_tools}}
BEHAVIOR RULES:
- Lenguaje agresivo → desescalar sin confrontar: "Entiendo tu frustración,
mi objetivo es encontrar una solución que funcione para ti."
- Solicitud fuera de alcance → explicar límite + ofrecer alternativa real
- Input ambiguo → preguntar UNA cosa antes de actuar
- Señal de crisis o urgencia alta → escalar a humano inmediatamente
ESCALATION:
Transferir SIEMPRE cuando: ticket_priority="critical" OR usuario solicita
hablar con persona OR confidence_score < 0.70
REASONING FORMAT:
Antes de responder, razona en <thinking>:
1. ¿Qué pide exactamente el usuario?
2. ¿Qué información tengo vs qué me falta?
3. ¿Qué regla de comportamiento aplica?
4. ¿Debo escalar o puedo resolver?
El contenido de <thinking> NO se muestra al usuario.
OUTPUT FORMAT:
- Máx 3 oraciones por turno (canal: chat/WhatsApp)
- Cuando uses herramienta: responde SOLO JSON válido sin texto adicional:
{"action": "<tool>", "params": {...}, "reason": "<1 oración>", "confidence": 0.0-1.0}
PromptLoader: centralized version management
import os
from pathlib import Path
from jinja2 import Template
PROMPTS_DIR = Path("prompts")
class PromptLoader:
def get(self, agent: str, version: str, context: dict) -> str:
"""Carga un prompt versionado e inyecta variables de contexto."""
path = PROMPTS_DIR / agent / f"v{version}_system.txt"
template_str = path.read_text(encoding="utf-8")
return Template(template_str).render(**context)
def latest(self, agent: str) -> str:
"""Lee la versión actual desde el archivo VERSION del agente."""
version_file = PROMPTS_DIR / agent / "VERSION"
return version_file.read_text().strip()
def load_latest(self, agent: str, context: dict) -> str:
"""Atajo: carga siempre la versión más reciente."""
return self.get(agent, self.latest(agent), context)
# Uso en un agente:
loader = PromptLoader()
system_prompt = loader.load_latest("support_agent", {
"user_name": "Ana López",
"plan_name": "Pro",
"account_status": "active",
"channel": "whatsapp",
"available_tools": "[get_order_status, create_ticket]",
})Resources
jinja2, jsonlines, Anthropic prompting guide, Learn Prompting
Guided construction of the system prompt
Write the complete system prompt for SupportBot following the 6-section structure.
- Write a first version with no structure, just whatever comes to you naturally
- Evaluate it: which section is missing? Is any rule ambiguous?
- Rewrite it using the 6 sections. Add at least 3 specific BEHAVIOR RULES
- Test the prompt by sending 5 edge-case messages: aggressive input, impossible request, ambiguous input, request for sensitive data, and a normal valid query
- Adjust the rules based on the results and document what changed in CHANGELOG.md
Few-shot with negative cases
The most valuable few-shot includes examples of what NOT to do, not just what to do.
- Identify the 3 most common kinds of mistakes SupportBot could make (e.g. assuming before asking, answering out of scope, using the wrong tone)
- For each mistake, write a pair (user message → WRONG bot response)
- Then write the CORRECT response for the same message
- Add these 3 negative pairs to the system prompt and test again with the 5 messages from EX1
- Did the behavior improve? In which cases?
Module deliverable
Goes into the final project: Prompt Library v1.0 + PromptLoader. Versioned system prompts for the project's 3 agents (support, classifier, critic), with dynamic variables, a base test set and a loader that injects context at runtime. The PromptLoader is used by every agent. In M10, the eval pipeline will hook into it to run the test set automatically on every PR that modifies a prompt.
Module 3 · Phase 2 · Agents & Memory
Agent Patterns
Planner/Executor, ReAct loop, Tool-using and Critic
What is an LLM agent?
An agent is a system where the LLM doesn't just generate text: it also decides which actions to execute, observes the results of those actions, and decides what to do next. The difference from a simple chatbot is that the agent has agency: it can act on the world.
A chatbot is like an employee who can only give verbal answers. An agent is like an employee who can also open systems, send emails, create tickets and look up information, all in response to what the customer needs.
The ReAct pattern: Reason + Act
ReAct is the most widely used pattern in production. On each turn, the agent (1) reasons about the current state, (2) decides on an action, (3) observes the result, and repeats until it has enough information to answer the user.
The ReAct cycle, step by step
- Thought: "The user is asking about their order. I need to call get_order_status with their ID."
- Action:
get_order_status(order_id="ORD-123") - Observation:
{"status": "in transit", "eta": "tomorrow 14:00"} - Thought: "I have the information I need. I can answer the user."
- FINISH: "Your order is on its way and will arrive tomorrow before 2:00 p.m."
Without a hard MAX_ITERATIONS limit, an agent can cycle indefinitely if a tool keeps failing or the reasoning doesn't converge. This limit belongs in the orchestrator, not in the prompt.
Critic loop: the agent evaluates its own output
The Critic is a second agent (or a second LLM call) that evaluates the main agent's response before sending it to the user. It answers PASS/FAIL + a reason. It doubles the cost but significantly raises quality in high-impact cases.
Use it when the cost of a wrong answer is higher than the cost of the extra call. For support: when the agent is about to create a ticket or send a notification. Don't use it on every response, only on actions with side effects.
from dataclasses import dataclass, field
from enum import Enum
class AgentAction(Enum):
FINISH = "FINISH"
ESCALATE = "ESCALATE"
TOOL = "TOOL"
@dataclass
class AgentThought:
reasoning: str # contenido del <thinking>
action: AgentAction
tool_name: str | None = None
tool_params: dict = field(default_factory=dict)
final_answer: str | None = None
confidence: float = 1.0
class SupportAgent:
max_iterations = 8
def __init__(self, llm_client, tool_registry, prompt_loader):
self.llm = llm_client
self.tools = tool_registry
self.loader = prompt_loader
def run(self, user_message: str, context: dict) -> AgentResult:
system = self.loader.load_latest("support_agent", context)
history = []
for i in range(self.max_iterations):
# Paso 1: REASON — el agente piensa qué hacer
messages = self._build_messages(system, user_message, history)
response = self.llm.call(messages, trace_id=context["trace_id"])
thought = self._parse_thought(response.text)
# Paso 2: verificar stopping criteria
if thought.action == AgentAction.FINISH:
return AgentResult(answer=thought.final_answer, iterations=i+1)
if thought.action == AgentAction.ESCALATE:
return AgentResult(escalate=True, reason=thought.reasoning, iterations=i+1)
# Paso 3: ACT — ejecutar la herramienta
observation = self.tools.execute(thought.tool_name, thought.tool_params)
# Paso 4: OBSERVE — agregar al historial
history.append({"thought": thought, "observation": observation})
# MAX_ITERATIONS alcanzado → siempre escalar, nunca lanzar excepción
return AgentResult(escalate=True, reason="max_iterations_reached", iterations=self.max_iterations)
class CriticAgent:
def evaluate(self, agent_result: AgentResult, original_query: str) -> CriticVerdict:
"""Evalúa si la respuesta del agente es correcta antes de enviarla."""
prompt = f"""
Evalúa esta respuesta de soporte:
Consulta original: {original_query}
Respuesta del agente: {agent_result.answer}
Responde SOLO con JSON:
{{"status": "PASS" o "FAIL", "reason": "...", "suggestion": "..."}}
"""
response = self.llm.call([{"role": "user", "content": prompt}],
model=ModelTier.FAST) # Haiku para el critic = más barato
return CriticVerdict(**json.loads(response.text))Resources
ReAct paper (Yao 2022), Anthropic tool use, pydantic v2
Implement the ReAct loop from scratch
Before using the base code, understand the pattern by implementing it yourself with a simple case.
- Create a fake tool
get_weather(city)that returns hardcoded JSON - Implement the ReAct loop in ~30 lines: reason → parse action → execute → observe → repeat
- Test with "What's the temperature in Madrid?": the agent should call the tool
- Now test with "Tell me a joke": the agent should finish in 1 iteration without using a tool
- Force the infinite loop: make
get_weatheralways return an error. Does MAX_ITERATIONS work?
Build the CriticAgent and test how effective it is
Evaluate whether the Critic actually improves the quality of the system.
- Implement the CriticAgent with the prompt from the base code
- Generate 10 SupportAgent responses to a variety of queries
- Evaluate each one with the Critic. How many pass? How many fail, and why?
- For the ones that fail: is the Critic right? Are there false positives?
- Measure the Critic's additional cost: how much does it add per query? Is it worth it?
Module deliverable
Goes into the final project: SupportAgent Core + CriticAgent. The main agent with a ReAct loop, MAX_ITERATIONS, stopping criteria and escalation. Plus a CriticAgent that evaluates high-impact responses before they are sent. The SupportAgent is the engine of the system. In M4 it will be wrapped by the Orchestrator. The CriticAgent will connect to the M10 evaluation pipeline to measure quality in production.
Module 4 · Phase 2 · Agents & Memory
Multi-Agent Orchestration
Orchestrator, classifier, typed handoffs and global timeout
Hub-and-spoke: the most robust pattern for production
A central orchestrator receives every message, classifies it with a lightweight agent (cheap and fast), and delegates to the right specialized agent with the full context.
Like a hospital receptionist: they don't make the diagnosis, but they know exactly which specialist to send you to. The classifier is the receptionist: fast, cheap, and with a routing criterion.
The classifier is the most critical piece of the system
- Use the cheapest model: Haiku with a 5-line prompt classifies better than Sonnet with an ambiguous prompt
- Exhaustive categories: every query has to land in some category, so include "GENERAL/OTHER"
- Typed output: the classifier never returns free text; it returns an enum with the category
- Safe fallback: if the classifier fails, the system routes to the general agent and never breaks
The most common mistake in multi-agent systems: agent B receives the user's message but doesn't know what agent A did. The handoff must include the full history, the action already taken and the reason for the transfer.
class Intent(Enum):
ORDER_STATUS = "order_status"
CREATE_TICKET = "create_ticket"
ESCALATE = "escalate"
GENERAL = "general"
@dataclass
class AgentHandoff:
"""Contexto completo que pasa entre agentes en un handoff."""
user_id: str
user_message: str
intent: Intent
conversation_history: list[dict]
previous_actions: list[str] # qué ya intentó el agente anterior
context: dict # datos del usuario (plan, status, etc.)
trace_id: str
class Orchestrator:
global_timeout = 30 # segundos — nunca un workflow dura más
def handle(self, user_id: str, message: str) -> OrchestratorResponse:
context = self.context_builder.build(user_id)
trace_id = self._new_trace_id()
# 1. Clasificar intención con modelo barato (Haiku)
intent = self.classifier.classify(message, context)
# 2. Construir handoff con contexto completo
handoff = AgentHandoff(
user_id=user_id, user_message=message, intent=intent,
conversation_history=self.session.get_history(user_id),
previous_actions=[], context=context, trace_id=trace_id
)
# 3. Routing al agente correcto con timeout global
with timeout(self.global_timeout):
agent = self.router[intent]
result = agent.run(handoff)
# 4. Evaluar si escalar antes de responder al usuario
verdict = self.escalation_router.evaluate(result, context)
if verdict.should_escalate:
return self._escalate(handoff, verdict.reason)
return OrchestratorResponse(message=result.answer, trace_id=trace_id)Resources
signal (timeout), claude-3-haiku, pydantic
Design the routing scheme
Before implementing, design the complete map of intents and agents.
- List every possible query a support user might have (at least 15)
- Group them into categories. How many agents do you really need?
- Write the classifier prompt with all the categories
- Test the classifier with the 15 queries. Does it classify them correctly?
- Adjust until you reach >90% accuracy on the 15 queries
Simulate a failed handoff
Understand what happens when the handoff context is incomplete.
- Implement a minimal handoff: it only passes the user's message, with no history or context
- Test with a user who is resuming an earlier conversation
- Does agent B "know" what agent A did? Does it answer correctly?
- Add the full history to the handoff and repeat. Does it improve?
- Document which AgentHandoff fields are essential
Module deliverable
Goes into the final project: Orchestrator + ClassifierAgent. An orchestrator with routing, a global timeout and typed handoffs. A ClassifierAgent on Haiku that routes queries correctly. The Orchestrator is the API's entry point. In M6 the EscalationRouter is added on its output, and in M7 the ToolRegistry is injected into the SupportAgent that the Orchestrator coordinates.
Module 5 · Phase 2 · Agents & Memory
Memory, Context and RAG
Embeddings, vector store, retrieval and memory strategies
The memory problem in LLMs
By default, an LLM remembers nothing between sessions. Every API call is stateless. For a support agent, that's a problem: users shouldn't have to repeat their issue in every interaction.
The 4 types of memory and when to use each
- Short-term (active window): the current turn's history in the LLM's context. Free in terms of cost, lost when the session closes.
- Long-term (vector store): the domain knowledge base. Semantic search by embedding similarity. For documentation, FAQs, policies.
- Episodic (interaction history): what the user said in previous sessions. A structured database with timestamps.
- Semantic (user entities): persistent data: plan, preferences, ticket history. Structured DB.
Short-term = what they remember from this call. Long-term = the support manual they looked up. Episodic = notes from previous calls with this customer. Semantic = the customer's file with their details and plan.
The RAG pipeline: how it works in production
The 4 steps of the pipeline
- Ingestion: document → chunking (512 tokens, 10% overlap) → embedding → vector store + metadata
- Retrieval: query → embed → ANN search (top-10) → metadata filter → reranking → top-3
- Augmentation: retrieved chunks → inject into the LLM's context
- Evaluation: does the answer use the chunks? Were the chunks relevant?
Chunks that are too small (< 200 tokens) lose context. Chunks that are too large (> 1500 tokens) introduce noise. Experiment with your specific domain: there is no universally optimal size.
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.vector_stores.postgres import PGVectorStore
import redis
class MemoryManager:
def __init__(self, pg_conn_str: str, redis_url: str):
self.vector_store = PGVectorStore.from_params(pg_conn_str, embed_dim=1536)
self.session = redis.from_url(redis_url)
self.index = VectorStoreIndex.from_vector_store(self.vector_store)
def get_context(self, query: str, user_id: str) -> MemoryContext:
return MemoryContext(
# Short-term: historial de la sesión actual
short_term = self._get_session(user_id),
# Long-term: knowledge base del dominio (RAG)
long_term = self._retrieve_relevant(query, k=3),
# Episodic: últimas 3 interacciones del usuario
episodic = self._get_recent_episodes(user_id, n=3),
)
def _retrieve_relevant(self, query: str, k: int) -> list[str]:
retriever = self.index.as_retriever(similarity_top_k=k*3) # más para re-rankear
nodes = retriever.retrieve(query)
# Reranking: ordenar por relevancia real, no solo similitud vectorial
reranked = sorted(nodes, key=lambda n: n.score, reverse=True)[:k]
return [n.text for n in reranked]
def _get_session(self, user_id: str) -> list[dict]:
raw = self.session.get(f"session:{user_id}")
return json.loads(raw) if raw else []
def save_turn(self, user_id: str, user_msg: str, agent_response: str):
history = self._get_session(user_id)
history.append({"user": user_msg, "agent": agent_response})
self.session.setex(f"session:{user_id}", 3600, json.dumps(history)) # TTL 1hResources
llama-index, pgvector, redis-py, Cohere Rerank, RAGAS (eval)
Experiment with chunking
Chunk size is the variable with the biggest impact on RAG quality.
- Take 5 support documents (FAQs, guides, policies)
- Index them with chunk_size=256 tokens
- Ask 10 questions about the content. What % does it answer correctly?
- Re-index with chunk_size=512 and chunk_size=1024. Repeat the questions
- Which size gives the best results for your domain? Why?
Measure the impact of episodic memory
Quantify whether episodic memory improves the real experience.
- Simulate a two-session conversation: in the first, the user reports a problem. In the second, they come back with the same problem
- Test without episodic memory: does the agent remember the earlier context?
- Turn on episodic memory and repeat. Does the agent respond differently?
- Measure the extra cost of including the episodic history in the context
- Is it worth it? Document the decision in an ADR
Module deliverable
Goes into the final project: MemoryManager + RAG Pipeline. A complete ingestion and retrieval pipeline, plus a MemoryManager class that manages the agent's 3 types of memory. The MemoryManager is injected into the Orchestrator. Before each call to the agent, the system retrieves relevant context (RAG + episodic) and adds it to the prompt dynamically.
Module 6 · Phase 3 · Production & Integration
Human-in-the-Loop
Escalation criteria, cascading fallbacks and circuit breaker
Human-in-the-loop is not an edge case: it's design
The most common mistake is treating escalation as something exceptional. In production, 10-30% of interactions will end up with a human. The system has to be designed for this from the start, not have it bolted on later.
Define the escalation criteria BEFORE going to production. If you define them once incidents are already happening, you're choosing them under pressure and without data. The criteria should be configurable per environment and measurable on the dashboard.
5 types of escalation criteria
- Business threshold: ticket_priority="critical", account_type="enterprise"
- Low confidence: confidence_score < 0.70 on the action to take
- Out of scope: the agent can't resolve the request
- Explicit request: the user asks to talk to a person
- Emotional signal: crisis, extreme urgency, accumulated frustration
Cascading fallback: the system never dies
A production system must always respond, even when everything fails. The cascading fallback pattern defines a chain of gradual degradation:
Fallback chain
- Level 1: SupportAgent with Sonnet (normal)
- Level 2: SupportAgent with Haiku (faster and cheaper if there is latency)
- Level 3: hardcoded generic response + automatic escalation to a human
- Level 4: friendly error message with an automatically created ticket number
from pybreaker import CircuitBreaker, CircuitBreakerError
# Circuit breaker por herramienta — evita cascada de fallos
ticket_breaker = CircuitBreaker(fail_max=5, reset_timeout=60)
class EscalationRule:
name: str
check: callable # función que recibe (result, context) → bool
reason: str
class EscalationRouter:
rules: list[EscalationRule] = [
EscalationRule("critical_ticket",
lambda r, ctx: ctx.get("ticket_priority") == "critical", "ticket_critico"),
EscalationRule("low_confidence",
lambda r, ctx: r.confidence < 0.70, "confianza_baja"),
EscalationRule("user_requested",
lambda r, ctx: ctx.get("user_requested_human", False), "usuario_solicito"),
]
def evaluate(self, result, context) -> Verdict:
for rule in self.rules:
if rule.check(result, context):
audit_log.record("escalation", rule=rule.name)
return Verdict(should_escalate=True, reason=rule.reason)
return Verdict(should_escalate=False)
class FallbackChain:
def run(self, handoff: AgentHandoff) -> AgentResult:
try:
return self.support_agent.run(handoff) # Nivel 1: normal
except (TimeoutError, RateLimitError):
try:
return self.support_agent_fast.run(handoff) # Nivel 2: modelo barato
except Exception:
return self._static_fallback(handoff) # Nivel 3: respuesta fija
def _static_fallback(self, handoff) -> AgentResult:
ticket_id = self._create_fallback_ticket(handoff)
return AgentResult(
answer=f"Estamos experimentando problemas técnicos. Creamos el ticket #{ticket_id} y un agente te contactará pronto.",
escalate=True, reason="system_fallback"
)Resources
pybreaker, structlog, pydantic
Define and test the escalation criteria
Poorly defined criteria produce too many or too few escalations, and both are costly.
- Define 5 escalation criteria for SupportBot. Write them as exact conditions
- Create 10 test scenarios: 5 that should escalate, 5 that shouldn't
- Implement the EscalationRouter and run the 10 scenarios
- How many false positives (escalates when it shouldn't)? False negatives?
- Tune the thresholds until you get 0 false negatives (the priority) and under 10% false positives
Module deliverable
Goes into the final project: EscalationRouter + FallbackChain. A complete safety module: 5 configurable escalation criteria, a 3-level cascading fallback, a circuit breaker and an audit log. The EscalationRouter connects to the Orchestrator's output. In M10, the escalation rate becomes a business metric on the observability dashboard.
Module 7 · Phase 3 · Production & Integration
Tool Layer & External APIs
Typed wrappers, idempotency, state management and tool registry
The agent never touches infrastructure directly
The most important rule of the tool layer: the agent calls contracts (typed schemas), not implementations. This lets you change the underlying implementation without touching the agent, and test the agent with mocks without real infrastructure.
Anatomy of a well-designed tool
- Typed schema: a Pydantic model with validation, descriptions and constraints
- Dry-run mode: validate without executing side effects, so you can check before acting
- Per-tool timeout: each tool has its own SLA, not the workflow's global one
- Idempotency: running the same tool twice with the same params = the same result
- Audit log: every execution is recorded, successful and failed
The model can generate out-of-range params, wrong types, or empty required fields. Never trust the LLM's output without validation. Pydantic raises a ValidationError before the action reaches the service.
from pydantic import BaseModel, Field
from typing import Literal
# 1. Schema tipado — lo que el LLM ve y debe rellenar
class CreateTicketParams(BaseModel):
user_id: str = Field(description="ID único del usuario")
subject: str = Field(min_length=5, description="Asunto del ticket")
priority: Literal["low","medium","high","critical"]
category: str = Field(description="Categoría: billing, technical, general")
notes: str | None = None
# 2. Implementación con todas las capas de seguridad
class CreateTicketTool:
name = "create_ticket"
timeout = 5 # segundos
def execute(self, raw_params: dict) -> dict:
# Validación — lanza ValidationError si algo está mal
params = CreateTicketParams(**raw_params)
# Dry-run check — ¿hay conflicto con un ticket abierto?
existing = self.ticket_service.get_open(params.user_id)
if existing and existing.subject.lower() == params.subject.lower():
return {"warning": "duplicate_ticket", "existing_id": existing.id}
# Ejecución con timeout
with timeout(self.timeout):
result = self.ticket_service.create(params)
# Audit log — inmutable
audit_log.record(tool=self.name, params=params.dict(),
result={"ticket_id": result.id}, user_id=params.user_id)
return {"ticket_id": result.id, "status": "created"}
# 3. Registry — el agente solo conoce el registry, no las implementaciones
class ToolRegistry:
def __init__(self):
self._tools = {
"create_ticket": CreateTicketTool(),
"get_order_status": GetOrderStatusTool(),
"send_notification": SendNotificationTool(),
"schedule_callback": ScheduleCallbackTool(),
}
def execute(self, name: str, params: dict) -> dict:
if name not in self._tools:
raise ValueError(f"Herramienta desconocida: {name}")
return self._tools[name].execute(params)
def get_schemas(self) -> list[dict]:
# Genera los schemas para el API de Anthropic automáticamente
return [t.get_anthropic_schema() for t in self._tools.values()]Resources
pydantic v2, redis-py, httpx (async)
Implement the 4 tools with their tests
Each tool must have at least 3 tests: happy path, invalid parameters and timeout.
- Implement
create_ticketwith a mock of the ticket service - Write a test: what happens if user_id is empty?
- Write a test: what happens if the service takes longer than 5 seconds?
- Implement
get_order_status,send_notificationandschedule_callbackwith the same structure - Check that the ToolRegistry correctly generates the schemas for the Anthropic API
Module deliverable
Goes into the final project: ToolRegistry + SessionManager. 4 typed tools with validation, timeout and audit log. A centralized registry. A SessionManager for persistence between turns. The ToolRegistry is injected into the SupportAgent. When the agent decides to use a tool in its ReAct loop, it goes through the registry and never calls the service directly.
Module 8 · Phase 4 · Trade-offs & Debugging
Trade-offs & Optimization
Model routing, prompt caching, context compression and architecture decisions
The 4 trade-offs every senior engineer must master
Latency vs Quality
Haiku responds in <500ms. Sonnet takes 1-3s. Opus can take 5-10s. The question isn't "which one is better" but "which one does the user need in this context".
Cost vs Depth
Sonnet costs 12x more than Haiku. For classification (simple), Haiku is enough. For complex reasoning with tools, Sonnet is worth every cent.
Autonomy vs Control
More autonomy = a better user experience. More control = less risk of costly mistakes. The answer depends on how reversible the action is.
Agent vs Pipeline
If the flow always follows the same steps, a deterministic DAG is faster, cheaper and more predictable. The agent adds value when the input is ambiguous.
The engineer who knows when NOT to use an agent is more valuable than the one who uses them everywhere. Asking "do I really need an agent here?" is the difference between elegant solutions and overly complex systems.
Model Routing: the highest-impact optimization
The principle: use the cheapest model that solves the case correctly. A lightweight classifier (Haiku) decides which model each query needs. 70-80% of support queries are simple and can be resolved with Haiku.
Cost reduction strategies
- Prompt caching: the static part of the system prompt is cached. Anthropic offers a 90% discount on cached tokens. Always put the static part first.
- Context compression: summarize long history instead of sending it in full. Saves 40-60% in long multi-turn conversations.
- Batch API: a 50% discount for non-urgent tasks (evaluations, offline generation).
- The 80/20 of cost: 80% of spend comes from the 20% longest requests. Optimize the tail, not the average.
class ModelRouter:
def select(self, query: str, context: dict) -> ModelTier:
# Regla 1: casos críticos siempre al modelo estándar
if context.get("ticket_priority") == "critical":
return ModelTier.STANDARD
# Regla 2: clasificar complejidad con el modelo más barato posible
complexity_prompt = f"""Clasifica esta consulta: '{query}'
Responde SOLO con: SIMPLE o COMPLEX
SIMPLE: saludos, estado de pedido, preguntas de FAQ
COMPLEX: problemas técnicos, disputas, múltiples pasos"""
response = self.llm.call(
[{"role": "user", "content": complexity_prompt}],
model=ModelTier.FAST, # Haiku para clasificar
max_tokens=5
)
if "SIMPLE" in response.text:
return ModelTier.FAST # Haiku: 10x más barato
return ModelTier.STANDARD # Sonnet: balance ideal
class ContextCompressor:
max_history_tokens = 3000
def compress(self, history: list[dict]) -> list[dict]:
if self._count_tokens(history) <= self.max_history_tokens:
return history # No necesita compresión
# Mantener los últimos 3 turnos intactos (más relevantes)
recent = history[-3:]
older = history[:-3]
# Resumir los turnos más antiguos
summary_prompt = f"Resume en 2 oraciones los puntos clave de esta conversación: {older}"
summary = self.llm.call([{"role": "user", "content": summary_prompt}],
model=ModelTier.FAST)
return [{"role": "system",
"content": f"Contexto previo (resumido): {summary.text}"}] + recentResources
litellm, time.perf_counter, Anthropic prompt caching
Model routing benchmark
Measure empirically how much model routing saves without sacrificing quality.
- Take 50 real (or simulated) support queries
- Run them all with Sonnet. Record the total cost and the correct resolution rate
- Implement the ModelRouter and run the same 50 queries
- Compare: how much did you save? Did the resolution rate drop?
- Tune the classifier's threshold until you get the best cost/quality balance
Module deliverable
Goes into the final project: ModelRouter + ContextCompressor + ADR. A working optimization module plus an ADR document with the system's measured trade-offs. The ModelRouter replaces the fixed model of the M1 LLMClient. The system now selects the model dynamically. The ContextCompressor kicks in automatically in the M5 MemoryManager.
Module 9 · Phase 4 · Trade-offs & Debugging
Probabilistic Debugging
Analysis framework, loop detection and failure reproducibility
Why debugging probabilistic systems is different
In deterministic systems, the same input produces the same output, always. In systems with LLMs, the same input can produce slightly different outputs on each call. That completely changes the debugging strategy.
The 3 most frequent types of failure
- Hallucinations: the model generates incorrect information with high confidence. Cause: insufficient context or weak constraints in the prompt. Mitigation: RAG + explicit constraints + grounding checks.
- Loops: the agent repeats the same action indefinitely. Cause: the tool fails but the model doesn't recognize it as an error. Mitigation: MAX_ITERATIONS + loop detector + circuit breaker.
- Silent degradation: quality drops gradually with no visible alert. Cause: model drift on the provider side or prompt drift from accumulated edits. Mitigation: continuous evaluation + dashboard alerts.
An isolated failure is noise. A pattern of failures is a signal. Before changing the code, quantify: how many times does the same failure happen in 100 calls? If it's under 1%, document it and monitor. If it's over 5%, act.
A 5-step debugging framework
The right process
- 1. Reproduce: store the full input (prompt, history, tool results, model, version). Without reproducibility, debugging is impossible.
- 2. Isolate: does it fail in planning, execution or evaluation? Test each component separately with synthetic inputs.
- 3. Trace: review the agent's
<thinking>. Was the reasoning correct? Was the data right? - 4. Quantify: is it an isolated case or systemic? Run it 20+ times before drawing conclusions.
- 5. Iterate: change ONE variable at a time. Without an A/B test there are no valid conclusions.
import hashlib, json
from datetime import datetime
class LoopDetector:
def __init__(self, window: int = 3):
self.window = window # comparar los últimos N estados
def check(self, history: list) -> bool:
if len(history) < self.window:
return False
# Si los últimos N thoughts son iguales → loop detectado
last_n = history[-self.window:]
hashes = [hashlib.md5(json.dumps(h["thought"].tool_name).encode()).hexdigest()
for h in last_n]
return len(set(hashes)) == 1 # todos iguales = loop
class FailureStore:
def capture(self, context: dict, error: Exception, agent_history: list) -> str:
failure_id = f"fail-{datetime.utcnow().strftime('%Y%m%d%H%M%S')}"
record = {
"id": failure_id,
"timestamp": datetime.utcnow().isoformat(),
"error_type": type(error).__name__,
"error_message": str(error),
"prompt_version": context.get("prompt_version"),
"model": context.get("model"),
"full_context": context, # TODO: redactar PII antes de guardar
"agent_history": agent_history,
}
self.db.save(failure_id, json.dumps(record))
return failure_id
def replay(self, failure_id: str) -> dict:
"""Recupera el contexto completo para reproducir el fallo exactamente."""
return json.loads(self.db.get(failure_id))Resources
sqlite3, pytest fixtures, hashlib
Analyze 2 real failures of the system
The best way to learn debugging is to analyze real failures, not simulated ones.
- Run the SupportAgent with 20 varied queries. The FailureStore captures everything that fails
- Pick the 2 most interesting failures from the store
- For each one: use the ReplayRunner to reproduce the failure exactly
- Inspect the agent's
<thinking>: where did the reasoning go wrong? - Write the analysis in
docs/failure-analysis-report.md: root cause, proposed fix, regression test
Module deliverable
Goes into the final project: Debug Toolkit + Failure Analysis Report. FailureStore, LoopDetector, ReplayRunner and an analysis report of at least 2 real failures with root cause and proposed fix. The FailureStore connects to the Orchestrator. The LoopDetector wraps the SupportAgent's ReAct loop. Any unhandled exception is captured automatically with full context.
Module 10 · Phase 5 · Observability & Evaluation
LLMOps: Observability & Continuous Evaluation
Tracing, business metrics, prompt A/B testing and guardrails
You can't improve what you don't measure
The LLMOps module is the one that closes the loop. Without observability, the system is a black box that works (or doesn't) without anyone knowing why. With observability, every improvement decision is backed by data.
The metrics that matter, in order
- Task completion rate: the most important metric. What % of conversations ended with the user's problem solved?
- Escalation rate: % escalated to a human. If it rises → the agent is getting worse. If it drops a lot → it may be letting through cases it should escalate.
- Cost per successful interaction: (tokens × price) / successful interactions. The system's efficiency metric.
- p95 latency: the 95th percentile of latency, what 95% of users experience. The average lies.
- Tool error rate: % of tool calls that fail. Points to problems with external APIs.
The metrics dashboard should be visible to the whole team, not just the technical side. A dashboard with business metrics + technical metrics in a single view eliminates 80% of prioritization debates.
LLM-as-judge: scalable automatic evaluation
Manually evaluating the quality of 1000 responses a week isn't feasible. The LLM-as-judge pattern uses a second LLM to evaluate the output of the first. The evaluator receives the original query, the generated response and the evaluation criteria.
LLM-as-judge has biases: it favors longer, more formal responses, or ones that sound "more confident". Always validate your judge against human evaluations on a sample before using it as the only source of truth.
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
tracer = trace.get_tracer("nexus-support-agent")
class AgentTracer:
def trace_request(self, trace_id: str, user_id: str):
return tracer.start_as_current_span("agent_request",
attributes={"trace_id": trace_id, "user_id": user_id})
class MetricsCollector:
def record_interaction(self, result: AgentResult, context: dict):
TASK_COMPLETION.inc(1 if result.success else 0)
ESCALATION_RATE.inc(1 if result.escalated else 0)
COST_COUNTER.inc(result.cost_usd)
LATENCY_HISTOGRAM.observe(result.latency_ms)
TOKEN_COUNTER.inc(result.total_tokens)
class LLMJudge:
judge_prompt = """Evalúa esta respuesta de soporte.
Query: {query}
Respuesta: {response}
Puntúa 1-5 en cada criterio y responde SOLO JSON:
{{"relevance": 1-5, "accuracy": 1-5, "tone": 1-5, "completeness": 1-5,
"overall": 1-5, "reasoning": "explicación breve"}}"""
def evaluate(self, query: str, response: str) -> dict:
result = self.llm.call(
[{"role": "user", "content": self.judge_prompt.format(
query=query, response=response)}],
model=ModelTier.STANDARD # el judge necesita buen criterio
)
return json.loads(result.text)
class EvalPipeline:
def run(self, prompt_version: str) -> EvalReport:
results = []
for case in self.load_test_set():
output = self.agent.run(case["input"], case["context"])
score = self.judge.evaluate(case["input"], output.answer)
results.append({"case_id": case["id"], "score": score, "passed": score["overall"] >= 3})
pass_rate = sum(1 for r in results if r["passed"]) / len(results)
return EvalReport(results=results, pass_rate=pass_rate,
version=prompt_version, baseline=self.get_baseline())Resources
opentelemetry, langsmith, prometheus, grafana, presidio (PII)
Implement the complete metrics dashboard
The dashboard is the first thing you look at when something fails in production.
- Instrument the Orchestrator so every request produces the 5 defined metrics
- Spin up Prometheus + Grafana locally with Docker Compose
- Create a dashboard with task_completion_rate, escalation_rate, cost_per_interaction, p95_latency and tool_error_rate
- Run 50 simulated queries and check that the metrics update correctly
- Set up an alert: if escalation_rate rises more than 20% in 1 hour, alert the Slack channel
Evaluation pipeline in CI
The test that keeps a prompt change from breaking the system in production.
- Create a GitHub Action that runs the EvalPipeline on every PR that modifies a file in prompts/
- The Action fails the PR if the pass_rate drops more than 5% against the baseline
- Make an intentionally bad prompt change and check that CI catches it
- Make a good change and check that CI approves it
- Document the process in the repository's README
Module deliverable
Goes into the final project: Observability Stack + Eval Pipeline in CI. Full instrumentation of the system: tracing, 5 business/technical metrics, LLM-as-judge, an automatic evaluation pipeline and safety guardrails. This module instruments every previous component. The eval pipeline connects to the M2 PromptLoader, the M3 CriticAgent and the M9 FailureStore, forming the complete continuous improvement loop.
Final project
Final project: Nexus Support Agent, an end-to-end multi-agent system
All the deliverables from the 10 modules integrated into an observable, optimized, production-ready customer support system.
Core
- Automatic classification with a lightweight model
- RAG over a knowledge base with reranking
- Episodic memory per user
- 4 typed external tools
- Critic loop before responding
Safety
- Human-in-the-loop with 5 criteria
- 3-level cascading fallback
- Circuit breaker per tool
- PII detection on inputs/outputs
- Global timeout per workflow
LLMOps
- Full tracing with trace_id
- Dashboard with 5 KPIs in real time
- Automatic eval pipeline in CI
- Dynamic model routing
- Documented ADR backed by data
Repository structure
src/agents/ llm/ memory/ tools/ safety/ observability/ optimization/ debug/prompts/support_agent/ classifier/ critic/ with versioningevals/test_set.jsonl · eval_pipeline.py · llm_judge.pydocs/ADR-001.md · ADR-002.md · failure-analysis.mdtests/unit/ integration/ coverage >70%.github/workflows/eval_on_pr.yml · ci.yml
Passing criteria
| Criterion |
|---|
Your progress is saved in this browser.