About tokens
Token conservation is the practice of reducing unnecessary context and overly long responses so AI workflows stay fast and cost-aware without losing quality. Tokens are the small units of text an AI model reads and generates. A user prompt or tool result is broken into tokens before the model can work with it.
In practical terms, token usage comes from retrieved context and generated output. More tokens usually means higher cost and more latency.
Importance of token conservation
Token conservation is not about making prompts vague or withholding necessary context. It is about sending the right context at the right time. This becomes especially important when teams move from experimentation to a real implementation. Conserving tokens helps lower operating cost and keeps interactions more responsive.
Carbon MCP approach
Carbon MCP introduces several patterns to reduce avoidable token usage:
- The carbon-builder skill lazy-loads only the Carbon guidance needed for the current task, rather than injecting full guidance into every request.
- Tools use multi-step retrieval where each step returns only the information needed at that moment, rather than a large block of unrelated content.
- Prompt templates ask the model not to restate or summarize tool output once the needed context has been retrieved.
- Sample prompts ask for exact files and a clear stop condition, which reduces unnecessary narration and extra turns.
- Structured retrieval helps the model avoid repeat searches and pulling in unrelated content.
Low token usage
- Prompt: "Build a Carbon React v11 page with a UI Shell header and a three-tile grid. Output only App.jsx and styles.scss."
- What happens: The AI retrieves only what it needs, stops when the named files are done, and does not restate tool results.
For more guidance on writing efficient prompts, see Prompts.