Applied AI Agents: Production RAG, Tool Calling & Multi-Agent Orchestration
Course curriculum
Deep dive into autoregressive transformer models (GPT-4, Gemini 1.5/2.0, Claude 3.5, and open-weight models like Qwen/Llama). Understand tokenization algorithms (Byte-Pair Encoding - BPE) and token-to-word ratios. Master LLM inference hyperparameters: Temperature (controlling probability entropy), Top-P (nucleus sampling), Frequency and Presence Penalties. Calculate operational cost and latency based on prompt and completion token counts.
Move from ad-hoc prompting to robust prompt engineering. Master structural role formatting (System, User, Assistant). Implement Chain-of-Thought (CoT) reasoning to force models to 'think' step-by-step before answering. Construct Few-Shot exemplars to anchor output formatting and domain nuance. Implement prompt injection defenses: isolating untrusted user text inside XML/Markdown tags and instructing models to ignore embedded instructions.
Eliminate brittle regex parsing and hallucinated syntax. Force LLMs to generate 100% valid, type-safe JSON adhering strictly to JSON Schema definitions using native Structured Outputs (OpenAI, Gemini) and Pydantic models. Validate field types, optionality, string patterns, and value ranges. Build self-healing schema retry wrappers that pass validation error traces back to the model upon parsing failure.
Explore the mathematics of vector embeddings. Contrast sparse keyword vectors (TF-IDF, BM25) with dense semantic embeddings (OpenAI text-embedding-3, HuggingFace sentence-transformers). Understand how transformer encoders map textual semantics into high-dimensional geometric spaces (384 to 3072 dimensions). Calculate geometric distance metrics: Cosine Similarity, Dot Product, and Euclidean (L2) distance, and understand when each metric is appropriate.
Examine the architecture of dedicated vector storage engines: FAISS (Facebook AI Similarity Search), ChromaDB, Qdrant, and pgvector. Understand why exact k-Nearest Neighbors (k-NN) fails at scale. Dive into Approximate Nearest Neighbor (ANN) indexing structures: Inverted File Indexing (IVF) and Hierarchical Navigable Small World (HNSW) graphs. Benchmark trade-offs between search latency, index build time, and recall precision.
Build the document ingestion pipeline. Extract raw text from diverse formats: PDFs (using pdfplumber and PyPDF2), Markdown, and HTML. Compare text chunking strategies: naive fixed-character chunking, recursive character chunking based on semantic boundaries (paragraphs, sentences), and document-structure chunking. Configure chunk overlap to preserve context across boundaries and attach rich metadata (source, page number, timestamp) for filtered retrieval.
Diagnose the core failure modes of basic naive RAG: irrelevant retrieval, context fragmentation, hallucination despite retrieval, and context dilution ('lost in the middle'). Architect end-to-end production RAG pipelines featuring query pre-processing, multi-vector indexing, contextual compression, and strict hallucination guardrails in the synthesis prompt.
Maximize retrieval precision. Implement query transformation techniques: Sub-Question Decomposition (breaking complex queries into distinct sub-queries) and Hypothetical Document Embeddings (HyDE). Combine dense semantic vector search with sparse keyword search (BM25) using Reciprocal Rank Fusion (RRF). Apply Cross-Encoder re-rankers (Cohere Re-rank, BGE-Reranker) to score document-query relevance with full self-attention.
Optimize context window efficiency and auditable accuracy. Implement contextual compression to extract only the specific sentences relevant to the query from long retrieved chunks. Prompt the LLM to generate strict inline citations referencing specific document identifiers and page numbers. Build post-generation citation verification modules that verify every claimed quote exists verbatim in retrieved sources.
Turn passive LLMs into active agents. Master the function calling protocol (OpenAI, Gemini). Define tool declarations using standard JSON Schema (tool name, description, parameter types, enums, required fields). Understand how the model decides when to return a standard text completion versus a structured tool call request. Parse tool call arguments, execute native code, and format tool results into tool-role messages.
Equip agents with real-world utility. Build production-grade tool integrations: live web search (DuckDuckGo, Tavily), read-only SQL query execution with automatic schema introspection, and external REST API consumers. Enforce strict safety guardrails on database tools: enforcing SELECT statements only, read-only database connections, statement execution timeouts, and result set row limits to protect server memory.
Build resilient tool execution engines. Handle common real-world failures: models calling nonexistent functions, passing invalid arguments that fail Pydantic validation, or tools throwing network exceptions. Instead of crashing, capture exception tracebacks, format them into helpful natural language error messages, and feed them back to the agent so it can self-correct. Sandbox arbitrary code execution tools in isolated environments.
Implement the classic autonomous agent pattern: ReAct. Explore the cognitive loop: Thought (verbal reasoning over current state) -> Action (selecting a tool) -> Action Input (providing tool parameters) -> Observation (ingesting tool output) -> Reflection -> Final Answer. Build a clean, dependency-free ReAct execution engine in pure Python. Enforce iteration limits to prevent infinite execution loops and runaway API costs.
Overcome myopic decision-making in complex multi-step objectives. Understand why step-by-step reaction often wanders off-track. Implement the Plan-and-Solve architecture: (1) Planner Agent decomposes a complex objective into an explicit list of sub-tasks, (2) Executor Agent iterates through the plan executing actions, (3) Re-planner updates remaining tasks based on intermediate findings. Contrast linear execution with tree-of-thought exploration.
Endow agents with stateful persistence. Architect multi-tiered memory systems: Working Memory (current execution stack and tool observations), Short-Term Conversational Memory (sliding context buffers with automated LLM summarization of older dialogue turns), and Long-Term Episodic Memory (persisting past goals, user preferences, and solutions in vector databases for semantic retrieval across sessions).
Scale single agents into specialized agent teams. Explore collaboration topologies: Supervisor-Worker (central router agent delegates subtasks to specialized researcher, writer, and coder agents), Sequential Handoff (assembly line pipeline), and Multi-Agent Debate (opposing agents critique and refine outputs to minimize hallucination). Define explicit agent communication schemas and data contracts.
Tame non-deterministic LLMs with deterministic execution graphs. Model agent workflows as stateful graphs: State schemas (shared TypedDict context), Nodes (agent actions or tool invocations), and Conditional Edges (routing logic based on output evaluation). Implement cyclic loops (e.g. Code -> Test -> Fail -> Re-code -> Test -> Pass). Incorporate Human-in-the-Loop checkpoints to gate sensitive actions (e.g. sending emails or issuing payments).
Deploy AI agents safely in enterprise settings. Implement multi-layered guardrails (NeMo Guardrails, Llama Guard): PII redaction (masking credit cards, phone numbers, emails), topic boundary enforcement (preventing agents from answering off-topic queries), and hallucination checks. Implement operational cost governance: tracking token expenditures per user, per-session rate limits, and hard budget circuit breakers.
Establish rigorous quality benchmarks for stochastic AI agents. Move beyond manual inspection. Implement automated evaluation frameworks (Ragas, DeepEval). Master core quantitative metrics: Faithfulness (hallucination detection against retrieved context), Answer Relevance, Context Precision, and Tool Calling Accuracy. Set up automated regression test suites that run in CI/CD before deploying agent updates.
Monitor multi-step agent reasoning in production. Instrument agent applications with distributed tracing (LangSmith, Phoenix/Arize, OpenTelemetry). Trace execution spans across nested tool calls, retrievals, and LLM calls. Measure token throughput and Time-to-First-Token (TTFT). Implement semantic caching (GPTCache, Redis) to serve instantaneous answers to recurring semantic queries, slashing API latency and cost.
Package and deploy autonomous agent applications. Build high-performance asynchronous web APIs using FastAPI. Implement Server-Sent Events (SSE) to stream real-time reasoning steps, tool execution statuses, and generated answer tokens to web clients. Containerize the application using Docker, secure API keys in production .env files, and deploy behind an Nginx reverse proxy on an AWS EC2 instance.