The stack my agents run on.
A curated list of agent tooling, grouped by what it does. Mostly harnesses. Orange means it's running in production.
Harnesses & agent CLIs 13
Claude Code anthropic.com How I build now. Agentic coding, most of this site included. In use Pi pi.dev Minimal open-source agent CLI. Four tools, model-agnostic, self-extends at runtime. In use OpenCode opencode.ai Provider-agnostic terminal agent. The open Claude Code, basically. In use Hermes nousresearch.com Nous Research's local agent. Persistent memory, six terminal backends incl. Modal. In use Codex CLI github.com OpenAI's terminal agent. The other side's Claude Code. In use Aider aider.chat Git-native pair programmer in the terminal. The one everyone benchmarks against. Watching Gemini CLI github.com Google's open terminal agent. Huge free tier, huge context window. Watching Goose block.github.io Block's open agent harness. Extensible, MCP-first. Watching Crush charm.land Charm's glamourous terminal agent. Pretty and provider-agnostic. Watching Amp ampcode.com Sourcegraph's agent. Built for big codebases and teams. Watching Cursor cursor.com Used it hard before Claude Code. Still the bar for inline edits. Watching LangGraph github.com Graph-based orchestration. One of the harnesses I keep comparing against my own. Watching OpenClaw openclaw.ai Steinberger's personal-agent gateway. Wires your chat apps to an agent. 100k stars in a week. Tried
Models & gateways 2
Workflow & infra 4
Temporal temporal.io Durable execution under my agents. State that survives restarts, retries and long waits. In use Cloudflare cloudflare.com Edge for the front of the stack. Pages, Workers, the dumb-fast layer. In use Modal modal.com Serverless GPU that doesn't make me think about infra. On the shortlist. Watching vLLM github.com If self-hosting open models ever pays off, this is the serving layer. Watching
Boundaries & quality 3
Pydantic pydantic.dev Typed contracts at the boundary. Structured output I can actually trust. In use Datadog datadoghq.com Where I watch agents in production. Traces, logs, cost, the lot. In use Braintrust braintrust.dev Evals as a first-class loop. Tried it on a real agent, liked the ergonomics. Tried
Nothing here for that filter.