A good AI agent system does not lock its intelligence to a single tool or stuff all its memories into one conversation.

Raymond’s explanation

When many people build AI agents, their first response is to switch models, install more MCPs, and connect more tools.

After going through that cycle himself, Raymond instead concluded that two things really matter: tool neutrality and Context Compact.

Tool neutrality means first defining “what capabilities the task needs,” then deciding which tool to use. Codex may be sufficient for pure text judgment, reading files, and organizing. Claude may be more suitable when MCPs, external services, or higher permissions are needed. Use an API only when a product feature or automation system needs to make calls over the long term. Browser control usually comes last because it is slowest and most susceptible to UI changes.

Context Compact means not putting everything into the same context. Keep only rules that are needed every time in the core instructions; put domain SOPs in Skills; reusable knowledge in Knowledge Cards; project status in one active document; and archive old plans as references. This way, the AI does not have to work while carrying an entire warehouse every time it starts.

These two things are really two sides of the same idea:

  • Tool neutrality prevents a system from being taken hostage by a single AI tool.
  • Context Compact prevents the AI from being slowed by too much background information.
  • Together, they turn an AI agent from “useful once” into something “maintainable over the long term.”

Core judgments

Before asking AI to do something, ask five questions:

  1. Is this a pure reasoning task, or does it require operating an external system?
  2. Does the task involve risks from writing data, deployment, batch changes, payment, or accounts?
  3. Is the needed information a current live state or stable knowledge?
  4. Will this tool consume too much context for only a small amount of convenience?
  5. After this conversation ends, is there anything worth consolidating into a Skill, Knowledge Card, or project document?

If these five questions cannot be answered clearly, the problem is usually not the model but a poorly designed system entry point.

Implementation principles

  • Start with a shared entry point: Use a provider-neutral wrapper such as call_agent() to unify tasks, permissions, timeouts, fallbacks, and event logs before deciding whether to call Codex, Claude, or an API underneath.
  • Separate permissions before tools: Read-only tasks can take a faster, lower-cost path; MCP, high-permission, and external write tasks should take a more conservative path.
  • Layer context: Keep core rules short, read Skills as needed, store concepts in Knowledge Cards, keep status in project documents, and retain history in archives.
  • Correct documents with live status: Do not trust old documents about production, health, deployment, or bot status. Check the actual runtime first, then update documents for people.
  • One task per session: Plan long tasks first and execute them in a separate session. Compact or hand off as the context limit approaches rather than pushing through.
  • Leave observability: AI calls should record provider, mode, status, latency, and fallback reason. Without records, it is difficult to know whether the system is improving.

Common pitfalls

  • Installing every MCP one sees, then having the tool list consume the context window.
  • Turning AGENTS.md, CLAUDE.md, or core instructions into an encyclopedia that forces AI to load much irrelevant material every time.
  • A document says “completed,” but production has not actually switched over or the switch was never verified.
  • Treating “model selection” as architecture, then scattering every workflow across different providers’ features.
  • Using browser automation too early for tasks that APIs, CLIs, or MCPs could handle.

Example: Kairos

The Discord “Ray Xiaomeng” agent in Kairos ultimately adopted a provider-neutral runtime: pure-text read-only work defaults to Codex-first; tasks needing MCPs, external services, or higher permissions go to Claude; Claude is the fallback if Codex fails; each call is written to an event log, and the Dashboard shows recent events and fallback status.

The point is not “Codex beats Claude” or “Claude beats Codex”; it is that the system now has an intermediary layer that can replace the provider. If tool capability, price, quota, or reliability changes, only the routing needs adjustment; the whole bot does not need to be dismantled and rewritten.

Angles for future content

  • “Don’t Lock Your AI Agent to One Tool: How I Designed a Dual Codex / Claude System”
  • “Context Compact: Why Your AI Assistant Gets Dumber the More You Use It”
  • “More MCPs Are Not Always Better: The Hidden Cost Is the Context Window”
  • “From Slimming Down CLAUDE.md to Skills: How an AI Counterpart Layers Memory”

Origin

  • Boris Cherny’s tips on using Claude Code: shared project context, CLAUDE.md layering, and tool feedback loops.
  • OpenClaw observations and notes: Context Fork, PreCompact, multi-agent routing, and observations on the security model.
  • Tools Extend Thinking: Tools shape how people think; they are not just accelerators.
  • Kairos implementation of the provider-neutral runtime and Dashboard event log on 2026-06-15.

Tools Extend Thinking, AI Multimedia Production Toolchain, Automation, Second Brain