For years "AI coding" meant autocomplete: you typed, the tool suggested. Agentic CLI coding tools changed the deal. Instead of completing a line, they take a task in plain English — "find and fix the failing tests," "refactor this module," "add a login page" — then plan, read files, run commands, write code, and iterate until the job is done, all inside your terminal. This guide shows how to set one up and build your first agentic workflow.
What an agentic CLI coding tool actually is
A coding agent is a loop, not a completion. The CLI tool sends your prompt plus the current state of your project to a large language model, the model decides which tools to call (read a file, run tests, edit code, search the repo), the harness executes those calls in your terminal, and the results feed back into the next model call. That loop repeats until the task is finished or the agent needs your input.
The key difference from an IDE copilot is autonomy plus shell access. The agent can run your test suite, inspect the failure output, change three files, re-run the tests, and show you the final diff. You stay in the loop by approving or rejecting its actions, but you stop babysitting every keystroke.
Step 1 — Pick the right tool for your setup
You do not need to build a harness from scratch; the ecosystem already has strong options. Pick based on your budget and model preferences:
- Claude Code (Anthropic): the official Anthropic CLI agent. Terminal-first, with IDE and web integrations, sub-agents for parallel work, and project instructions via
CLAUDE.mdfiles. Ships with paid Anthropic plans — Pro starts at $20/month (a June 2026 comparison verified this against vendor docs). Best when you want the tightest integration with Claude models. - Codex CLI (OpenAI): OpenAI's terminal agent for coding and terminal automation with sandboxing. Included in ChatGPT Plus ($20/month) per that same June 2026 comparison. Good fit if you already live in the OpenAI stack.
- Gemini CLI (Google): open-source agent built on Gemini. One June 2026 comparison found its free tier the most generous of any agentic tool — 1,000 requests per day on a personal Google account. Ideal for cost-sensitive developers and Google Cloud shops.
- Aider: terminal agent that is model-agnostic — you bring your own API key for whatever model you prefer. Free to install; you pay only your model provider. Git-native: it auto-commits each change so
git resetis your undo button. Closest competitor to Claude Code, per several 2026 comparisons. - Cline (formerly Claude Dev): open-source agent that lives in VS Code, with an explicit Plan/Act approval mode, built-in cost tracking, and a newer CLI 2.0 with headless mode for CI/CD. Works with OpenRouter, Anthropic, OpenAI, Gemini, and local models via Ollama.
If you want zero cost to start, install Gemini CLI (free tier) or Aider (free tool, pay only API usage). If you want maximum capability out of the box, start with Claude Code or Codex CLI.
Step 2 — Install and configure your agent
Installation is a one-liner in most cases. For Aider (the cheapest to experiment with):
pip install aider-chat
export ANTHROPIC_API_KEY=your-key-here
cd your-project
aiderFor Gemini CLI, the free tier means you can start without an API key at all:
npm install -g @google/gemini-cli
cd your-project
geminiFor Claude Code:
npm install -g @anthropic-ai/claude-code
cd your-project
claudeOnce inside, verify the basics work before trusting it with real code: ask it to summarize the repository structure, then to run the test suite. If it can read your project and execute pytest or npm test successfully, the agent loop is wired correctly.
Step 3 — Give the agent project context
Agents are only as good as the context they have. Every major tool supports a project instruction file that the agent reads at the start of each session:
- Claude Code:
CLAUDE.md - Cline:
.clinerules - Aider:
.aider.conf.yml
Keep it short and concrete: how to build and test the project, coding conventions, and anything that bites newcomers. A solid template:
# Project conventions
- Run tests with: npm test
- Lint with: npm run lint
- Never commit node_modules or .env files
- API routes live in /src/routes, components in /src/components
- Prefer small, focused commitsThis file compounds in value. Every time the agent does something you dislike — wrong test command, wrong folder — add the correction to the file. After a week, your agent behaves like a team member who has been onboarded properly.
Step 4 — Run your first autonomous task
Start with a bounded, verifiable task so you can judge the output easily. A good first task:
Run the test suite. If any tests fail, find the root cause,
fix it, and re-run the tests until they pass. Show me the diff.Watch what happens: the agent reads your instruction file, runs the tests, inspects failures, edits source files, re-runs the tests, and reports back. With Aider, each edit is a git commit you can review and undo with git reset. With Cline, each file edit and shell command waits for your approval in Plan/Act mode.
Then try a multi-file task — the kind that used to take an hour of manual work:
Add input validation to the user registration endpoint.
Return 400 with a clear error message for invalid emails
and passwords shorter than 12 characters. Add tests.A May 2026 benchmark roundup illustrates why tool choice matters here: the harness around the model can shift results by several points — same model, different agent wrapper, measurably different scores. SWE-bench Verified figures reported in May 2026 put Claude Code (on Opus 4.7) at 87.6% and Codex (GPT-5.5) at 82.7%, with Gemini CLI at 80.6% — but treat self-reported benchmarks as directional, not gospel.
Step 5 — Build a custom agent workflow
Once you trust the loop, the real power comes from composing agents into workflows. Three patterns are common in 2026:
The orchestrator pattern. Give one "planner" agent the task of breaking work into steps, then let sub-agents execute each step in parallel. Claude Code supports spawning sub-agents natively — one agent writes the migration while another updates the tests.
The review loop. Two agents, one role each: a builder writes the code, a reviewer critiques it. Point both at your linting rules and test suite, and iterate until both agree. This mirrors how senior engineers review pull requests.
The headless CI pattern. Run the agent in CI on a schedule: gemini -p "update dependencies and open a PR if tests pass" in a GitHub Action. Human review stays at the PR stage. Cline's CLI 2.0 headless mode and Gemini CLI's non-interactive flag are designed for exactly this.
Costs and safety guardrails
Agentic coding is not free computing. A single non-trivial task can burn 50k to 500k tokens, and one widely-cited 2026 breakdown notes Anthropic's own docs put average usage near $6 per developer per day for heavy frontier-model use. Set hard API spend caps before your first run, route routine work to cheaper efficient models, and keep the heavyweight models for genuinely hard problems.
Safety-wise, follow three rules. First, keep agents in a git repository so every change is reviewable and revertible. Second, restrict shell permissions — Codex CLI's sandboxing and Cline's approval-gated actions exist because an autonomous agent with full shell access is one bad instruction away from rm -rf on the wrong directory. Third, never let an agent near production credentials; use a sandboxed dev environment or Docker for autonomous runs.

