Free course lesson

Prompt Engineering Trade-offs | AI Agents from Scratch

The three levers every prompt trades: cost, latency, and accuracy

Course
AI Agents from Scratch
Lesson type
READING
Access
Free module

What this lesson covers

The three levers every prompt trades: cost, latency, and accuracy

A practical, ground-up introduction to AI agents. Understand the ReAct loop, tools, orchestration, memory, and evaluation, then build real agents using Aurora Trading Agent platform as your lab.

  1. Explain the difference between a bare language model, a chatbot, and an AI agent

    Learning outcome for this course.

  2. Describe the ReAct loop (Thought → Action → Observation) and why it makes agents agentic

    Learning outcome for this course.

  3. Design tool-calling workflows and understand why tools, function calling, and MCP servers are all the same core concept

    Learning outcome for this course.

  4. Use autonomy controls (whitelists, approvals) to safely run agents in production

    Learning outcome for this course.

  5. Connect memory, scheduling, and subagents into a full autonomous workflow

    Learning outcome for this course.

  6. Build traces and evaluators (including LLM judges) to measure agent quality

    Learning outcome for this course.

Lesson reading

Every time you write a system prompt, you're making trade-offs across three things: **how much it costs**, **how fast it is**, and **how accurate it is**. Pick any two - getting all three on a single prompt is rare, and usually means you're not pushing any of them.

This isn't theory. Every production AI app hits these trade-offs within its first week of real usage. This is how to think about them before you're surprised.

## Cost

Every API call has an **input token** bill and an **output token** bill. Input tokens include everything you send the model: system prompt, interpolated contexts, few-shot examples, conversation history, and the new user message. Output tokens are what the model generates back. Output tokens are typically **3–5× more expensive** than input tokens, which is why verbose models that "think out loud" can blow up your budget fast.

A few cost realities you'll hit:

- **Few-shot examples add tokens linearly.** Five examples of 200 tokens each = 1,000 input tokens on *every single call*. Multiply by thousands of users and you're paying for those examples forever. - **Context interpolation adds tokens.** That `${indicator}` in Generate Portfolios expands to hundreds of lines at runtime. Every call pays for it. - **Long system prompts cost on every request**, not just the hard ones. A 5,000-token system prompt on an easy request still costs 5,000 input tokens. - **Model choice dominates.** Gemini 3 Flash is roughly **20× cheaper** than Claude Sonnet for the same token count. Pick the right model before you try to shave 100 tokens off your prompt.

**Prompt caching** helps a lot. Anthropic and OpenAI both let you cache stable system prompts - you pay full price on the first call, then a fraction of that (often 10%) on every subsequent call that reuses the same prefix. If you're calling a big prompt thousands of times a day, enable this.

## Latency

Longer prompts take longer to process. The relationship is roughly linear: double the input tokens, expect the response to start about twice as slow. **Model choice matters more than prompt length for time-to-first-token** - Flash and Haiku models are typically 3–5× faster than their bigger siblings.

Other latency levers:

- **Streaming** makes long responses *feel* fast - the learner sees words appearing within a second, even if the full response takes 30. Use it for anything the user is reading in real time. - **Parallel tool calls** beat sequential. If your agent can call three tools that don't depend on each other, run them in parallel and merge results. - **Don't make the classifier expensive.** The router prompt runs on *every* user message. It has to be fast, cheap, and simple. NexusTrade's classifier is Gemini 2.0 Flash for exactly this reason - any slower and every message feels sluggish.

## Accuracy

Accuracy is where prompt engineering earns its name. The moves that actually work, in rough order of impact:

- **Zero-shot** (just instructions, no examples) is fine for simple, unambiguous tasks: "summarize this text", "classify as positive or negative". - **Few-shot** is a big jump for structured output. Three correct examples is the usual sweet spot; past 5–10, you hit diminishing returns and you're just paying more for marginal gains. - **Chain-of-thought** ("think step by step before answering") is a big jump for reasoning tasks - math, multi-step logic, debugging. It barely helps for lookups. - **Temperature 0** for anything deterministic. YAML, JSON, SQL - creativity is a liability when a parser is waiting. - **Structured output** (JSON mode, response schemas) eliminates a whole class of bugs by not giving the model the option to respond with prose. - **Watch system prompt length.** Models genuinely degrade as instructions get longer - typically around the 10k-token mark. If your prompt is bigger than that, you need to split it into subagents, not keep adding rules.

## How NexusTrade picks models

The hands-on exercises all use GPT 6 Luna because it's cheap and you're spending your own token budget. The real NexusTrade agent system picks different models for different jobs - these are the actual production choices:

- **The classifier** (routes every user message to one of 23 sub-prompts) → `google/gemini-2.5-flash-lite`. Cheap, fast, good at picking an integer from a list. Runs on *every* message, so it has to be near-free. - **Generate Portfolios** (turns natural language into NT SDK TypeScript strategies) → `openai/gpt-5.6-luna` at temperature 0. The NT SDK typechecks and validates the code — creativity is a liability when a compiler is waiting. - **Deep Research** (multi-source investment analysis reports) → Exa Agent (medium). Web-connected research with real source URLs — a dedicated research agent, not a chat model with search bolted on. - **Stock Screener** (generates SQL from natural language) → `anthropic/claude-haiku-5.5` with a schema-forced output. SQL is deterministic, the schema is finite, and we need it fast.

**The meta-pattern:** every specialized prompt deserves its own model. You pick based on which lever matters most for *that* task. The Stock Screener runs on every screener query, so cheap-and-fast matters. Deep Research runs once when the user explicitly asks for a big report, so paying for a dedicated research agent is fine. Understanding this is the real difference between a demo and a production AI app.

## The meta lesson

**Don't optimize prematurely.** Ship with a simple system prompt, measure cost-per-call and latency-per-call per prompt in production, then iterate on the ones that hurt. Prompt engineering is the only form of engineering where "just try it and see what works" is legitimately the right approach - the feedback loop is tight enough that you can test a change and see its impact within minutes.

But the moment you have real users, you need to start measuring. A prompt that costs $0.01 per call is fine until it's called 100,000 times a day, at which point it's $30,000 a year and you'll wish you had shaved half the examples. The difference between a product that scales and one that bleeds money is usually one person who was measuring.

## Further reading

- [Anthropic - Prompt Engineering Overview](https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/overview) - the canonical Claude guide, updated frequently - [OpenAI - Prompt Engineering Best Practices](https://platform.openai.com/docs/guides/prompt-engineering) - OpenAI's version, focused on GPT models - [Lilian Weng - Prompt Engineering](https://lilianweng.github.io/posts/2023-03-15-prompt-engineering/) - research-oriented deep dive with references to papers

Module 2: What Is an AI Agent?

Continue exploring