Local CLI · Python 3.12

CodeAgent

A local, CLI-based autonomous coding agent. The model reasons; deterministic software searches, executes, constrains, verifies, persists, and observes.

Capabilities

End-to-end agent mechanics with explicit safety boundaries and verifiable completion.

Tool-calling loop

Native provider tool calls with iteration, repeat, and budget guards.

Permission engine

Table-driven allow, ask, and deny before execution; single-use approvals.

Isolation

Git worktree by default, checkpoints, rollback, and optional Docker sandbox.

Verification

Checks pipeline and completion gate; no verified finish without evidence.

Durable runs

SQLite state, resume after interrupt or quota wait, full event trace.

Evaluation harness

40-task fixture suite with scripted replay in CI and optional live slices.

Demo

Terminal recording from a representative session on the payment service example.

Architecture

Dependencies point downward: the loop orchestrates; tools do not call back into it.

flowchart TB
  CLI[CLI_Typer_Rich]
  Loop[AgentLoop]
  Provider[Provider]
  Tools[ToolRegistry]
  Perm[PermissionEngine]
  State[StateStore]
  Iso[Isolation]
  Verify[Verification]
  Obs[Observability]

  CLI --> Loop
  Loop --> Provider
  Loop --> Tools
  Loop --> Perm
  Loop --> State
  Loop --> Iso
  Loop --> Verify
  Loop --> Obs
  Tools --> Perm
        

Full specification: DESIGN.md

Engineering

Python 3.12, strict mypy, Ruff, pytest on Ubuntu and macOS. 442 offline pytest cases in CI (live excluded), including a dedicated security suite.

Evaluation

Scripted replay validates the full loop in CI without API tokens; live slices are optional and quota-aware.

42 / 42 Scripted replay (CI)

40 tasks plus harness smoke through the real loop, tools, and permissions via FakeProvider.

442 Offline pytest cases

Including security tests; live provider tests are optional and excluded from the default check script.

40 tasks Fixture suite

Small repositories with check commands and reference transcripts for deterministic replay.

For model comparison under real API limits, run codeagent eval --live --slice core12 --resume locally. Daily Groq quotas cap throughput; runs resume across days. Treat live output as exploratory, not a published scorecard.

Baselines and reports: evals/baselines/.

Scope

Shipped through phase 8, tuned for small repositories and Groq free-tier context and quota limits.

Decision records · Build progress