Lean. Fast. Inexpensive.
A lean AI coding agent written in Rust. Every design decision optimizes exactly two things: API cost and speed.
A coding agent re-sends its entire scaffolding on every round. ragent keeps that scaffolding at roughly 1.3k input tokens — and puts a price on every token it spends.
Instructions are compressed to the point where nothing decorative survives a billing round.
agent.rs::SYSTEMedit_file sends exact matches, atomically — never a whole-file rewrite.
Lines, bytes and hit counts are bounded; truncation tells the model exactly how to page for more.
tools.rs constsA file tree plus symbol outlines — orientation without shipping a single file body.
repomap.rsStale tool output collapses to [elided], then middle history drops with tool-call orphan repair.
A price table prices cache hits separately (DeepSeek hits are ~50× cheaper) and applies peak-window multipliers.
config.rs::cost
Every run is appended to ~/.local/share/ragent/stats.tsv.
The TUI shows task $x and session $/✓ live;
headless mode prints the task cost plus a lifetime average. In the benchmark it was the
only agent whose self-reported cost matched the proxy-measured cost exactly.
A native binary starts in about five milliseconds and renders a terminal UI you can actually scroll. Nothing is loaded that isn't needed to answer the question.
Tool calls appear as they happen; text renders while the model is still talking.
api.rs::stream_chatread_file, grep, glob,
ls and repomap run concurrently on a
JoinSet.
A single-division character heuristic estimates tokens in O(n) with zero load time.
util.rs::est_tokensBlocking file ops move to spawn_blocking; target
and node_modules are never walked.
Only new lines are re-wrapped; only visible lines materialize per frame.
tui.rsTokio current_thread, thin LTO, codegen-units=1, stripped — and mold for relinks.
ragent, Pi, Codex CLI and Grok CLI all ran deepseek-flash
through a per-agent logging proxy. Usage comes from the provider's own billing field — not from
any agent's self-report.
| Agent | Tasks OK | Wall clock | API reqs | Input tokens | Output | Spend | $/✓ task |
|---|---|---|---|---|---|---|---|
| ragent | 8/8 | 76.6 s | 39 | 104,862 | 10,304 | $0.0090 | $0.00112 |
| pi | 8/8 | 150.7 s | 39 | 117,632 | 10,621 | $0.0096 | $0.00120 |
| grok | 8/8 | 201.8 s | 73 | 1,254,634 | 21,554 | $0.0295 | $0.00369 |
| codex | 6/8 | 1,382.4 s | 379 | 53,323,006 | 216,831 | $0.4199 | $0.07000 |
ragent and Pi tie on token economy; ragent is ~2× faster at identical request counts. Codex failed two medium tasks by timing out in a compaction loop — 173 iterations, 22.4M input tokens, $0.176 for one endpoint.
Measured request body → billed input on the first request of a fresh session. Difficulty scaled request count for lean agents; for heavy ones it scaled request size. ragent's per-request input never left the low thousands.
More context is not
more reliability.
The agent shipping two orders of magnitude more context was the only one to fail.
Terse schemas, precise semantics. File tools are confined to the project root;
bash is your shell.
read_file · offset/limit paging
edit_file · atomic multi-edit search/replace
write_file
bash · timeout, head+tail truncation
grep · regex, gitignore-aware, glob filter
glob
ls
repomap
Eleven source files, each with a single job. The run loop, the wire client, the context compactor, the repo scanner and the TUI are all separate, legible surfaces.
Full markdown rendering in assistant output — headings, fenced code with a gutter, lists with hanging indents, blockquotes, rules and boxed pipe tables — behind turn markers, a braille activity line and a segmented status bar.
Point it at OpenAI, DeepSeek, OpenRouter, Ollama, vLLM, LM Studio — anything that
speaks the chat-completions wire format. Configuration resolves CLI flag → environment →
~/.config/ragent/config.toml → defaults.
# build cargo build --release # configure export RAGENT_API_KEY=sk-… # or OPENAI_API_KEY ragent --init # write a commented example # run ./target/release/ragent # TUI ./target/release/ragent "fix the bug in src/api.rs" git diff | ./target/release/ragent "review this" # piped stdin
# ~/.config/ragent/config.toml api_key = "…" base_url = "https://api.deepseek.com/v1" model = "deepseek-flash" context_budget = 60000 # compaction above this max_output_tokens = 8192 bash_timeout_ms = 30000 price_input = 0.15 # $/1M for unknown models price_output = 0.60 price_cached = 0.003
Linux builds link with mold by default (apt install mold);
the x86_64 target is x86-64-v3. On macOS or Windows the config is a no-op.