Agent setup · Exercise 01
Models
Choose the right model for your agent and your budget.
Prompt for your agent
Copied! Now paste that into your agent...
Models
The model behind your agent is the single biggest factor in how useful it’ll be to you. More capable models write better code, catch more mistakes, and need less hand-holding, but they cost more per request. Cheaper or free models can be fine for small, well-defined tasks, but on anything nontrivial they often need more back-and-forth to get right, which can end up costing more in your time than a better model would have cost in money.
Subscription vs. pay-per-token
Most agent harnesses let you pay one of two ways: a flat monthly subscription tied to the model vendor’s own app (like a ChatGPT or Claude plan), or pay-per-token API billing. For everyday interactive coding, a subscription plan is usually the better starting point: it’s predictable, and it’s typically better value than paying per token for the same amount of use. Pay-per-token billing makes more sense for automation, CI, or workloads where a human isn’t in the loop deciding when to send the next message.
Free tiers exist on most harnesses, but they usually default to the vendor’s weakest available models. That’s fine for learning the tool itself, but if you find yourself fighting the model to get simple things right, that’s a sign it’s time to either add a paid subscription or switch to a stronger model on your current plan.
Context windows
Every model has a context window: the amount of text it can consider at once, measured in tokens (roughly ¾ of a word). This includes your messages, any files the agent has read, and its own previous responses. Flagship models today typically offer context windows from 100,000 tokens up to a million or more, which is enough for long sessions without much thought. Smaller or cheaper models often have much smaller windows (8,000-32,000 tokens), which can fill up quickly.
When a context window fills up, agents typically either summarize older parts of the conversation to make room, or start losing track of earlier instructions. If your agent starts behaving as if it’s forgotten something you told it earlier, that’s usually a context problem, not a reasoning problem. Keeping sessions focused on one task, and starting fresh sessions for unrelated work, goes a long way.
Cloudflare as a model provider
If you already have a Cloudflare account, two options are worth knowing about, though neither is required for this course:
- AI Gateway is a proxy in front of providers like OpenAI, Anthropic, and others. With Unified Billing, you can pay for multiple providers through one Cloudflare bill instead of managing a separate account with each one.
- Workers AI runs a curated set of open-source models directly on Cloudflare’s network.
Check your harness’s own documentation for how to point it at a custom provider or gateway.