ragent

Lean. Fast. Inexpensive.

A lean AI coding agent written in Rust. Every design decision optimizes exactly two things: API cost and speed.

5.6 MB binary ~5 ms startup ~4 MB RSS Apache-2.0 · v0.2.0
ragent — zsh — 96×28
┌─ ragent v0.2.0 ─────────────────────────────────────┐ │ model deepseek-flash │ │ root ~/lemon/agent │ │ context 60k budget · auto-compaction above │ │ lifetime 8/8 ✓ · $0.0011 avg / ✓ │ └─────────────────────────────────────────────────────┘ ❯ fix the cart-merge bug and add a wishlist test ◆ reading src/shop.py, src/cart.py, tests/test_shop.py … ◆ 4 read-only tools dispatched concurrently ● Patched cart.merge() to de-duplicate line items and added tests/test_wishlist.py (6 checks). [task ok] $0.0021 · ↑29.5k ↓2.2k ⛁84% · life $0.0011/✓ ❯
deepseek-flash 60k ctx task $0.0021 session $0.0007/✓ ▲ scrolled
8/8
tasks solved on the difficulty ladder
$0.0011
per successful task vs. the same model
76.6s
wall clock for the full eight-task suite
~62×
cheaper per solved task than the heaviest agent
Cost

Every word is billed.
So there are very few of them.

A coding agent re-sends its entire scaffolding on every round. ragent keeps that scaffolding at roughly 1.3k input tokens — and puts a price on every token it spends.

~120-token system prompt

Instructions are compressed to the point where nothing decorative survives a billing round.

agent.rs::SYSTEM

Search/replace edits

edit_file sends exact matches, atomically — never a whole-file rewrite.

tools.rs::edit_file_sync

Hard output caps

Lines, bytes and hit counts are bounded; truncation tells the model exactly how to page for more.

tools.rs consts

Repo map, injected once

A file tree plus symbol outlines — orientation without shipping a single file body.

repomap.rs

Two-stage compaction

Stale tool output collapses to [elided], then middle history drops with tool-call orphan repair.

ctx.rs::compact

A live dollar meter

A price table prices cache hits separately (DeepSeek hits are ~50× cheaper) and applies peak-window multipliers.

config.rs::cost
Accountability

Cost per successful task, not per token.

Every run is appended to ~/.local/share/ragent/stats.tsv. The TUI shows task $x and session $/✓ live; headless mode prints the task cost plus a lifetime average. In the benchmark it was the only agent whose self-reported cost matched the proxy-measured cost exactly.

ts model in out cached cost ok 1731… deepseek-flash 11242 1261 8192 $0.0012 ok 1731… deepseek-flash 13030 1179 10112 $0.0012 ok 1731… deepseek-flash 25657 2841 22528 $0.0022 ok 1731… deepseek-flash 29545 2157 24576 $0.0021 ok ─ lifetime ────────────────────────────────────── tasks 8/8 ✓ · avg $0.0011/✓ · total $0.0090
Speed

Speed you feel
in the first keystroke.

A native binary starts in about five milliseconds and renders a terminal UI you can actually scroll. Nothing is loaded that isn't needed to answer the question.

SSE, streamed per chunk

Tool calls appear as they happen; text renders while the model is still talking.

api.rs::stream_chat

Parallel read-only tools

read_file, grep, glob, ls and repomap run concurrently on a JoinSet.

agent.rs

No tokenizer dependency

A single-division character heuristic estimates tokens in O(n) with zero load time.

util.rs::est_tokens

Gitignore-aware walks

Blocking file ops move to spawn_blocking; target and node_modules are never walked.

tools.rs · repomap.rs

Incremental scrollback

Only new lines are re-wrapped; only visible lines materialize per frame.

tui.rs

A deliberately small build

Tokio current_thread, thin LTO, codegen-units=1, stripped — and mold for relinks.

Cargo.toml
Measured, not claimed

Four agents. One model.
One task suite.

ragent, Pi, Codex CLI and Grok CLI all ran deepseek-flash through a per-agent logging proxy. Usage comes from the provider's own billing field — not from any agent's self-report.

Eight tasks · wall clock, provider-billed tokens and spend · off-peak DeepSeek rates
AgentTasks OKWall clock API reqsInput tokensOutput Spend$/✓ task
ragent8/876.6 s39104,86210,304$0.0090$0.00112
pi8/8150.7 s39117,63210,621$0.0096$0.00120
grok8/8201.8 s731,254,63421,554$0.0295$0.00369
codex6/81,382.4 s37953,323,006216,831$0.4199$0.07000

ragent and Pi tie on token economy; ragent is ~2× faster at identical request counts. Codex failed two medium tasks by timing out in a compaction loop — 173 iterations, 22.4M input tokens, $0.176 for one endpoint.

The smoking gun

Measured request body → billed input on the first request of a fresh session. Difficulty scaled request count for lean agents; for heavy ones it scaled request size. ragent's per-request input never left the low thousands.

ragent~4 KB → ~1.3k tokens · 70–90% cached
pi~6 KB → ~1.6k tokens · 75–90% cached
grok~63 KB → ~16k tokens · ~93% cached
codex~650 KB → ~240k tokens · 98–100% cached

More context is not
more reliability.

The agent shipping two orders of magnitude more context was the only one to fail.

Tools

Eight tools.
No ceremony.

Terse schemas, precise semantics. File tools are confined to the project root; bash is your shell.

read_file · offset/limit paging edit_file · atomic multi-edit search/replace write_file bash · timeout, head+tail truncation grep · regex, gitignore-aware, glob filter glob ls repomap

Small enough to read in one sitting.

Eleven source files, each with a single job. The run loop, the wire client, the context compactor, the repo scanner and the TUI are all separate, legible surfaces.

src/api.rs OpenAI-compatible SSE client, retries src/tools.rs schemas + implementations, output caps src/ctx.rs token estimation + two-stage compaction src/repomap.rs repo skeleton builder src/agent.rs run loop, parallel dispatch, events src/tui.rs ratatui UI: scrollback, input, status bar src/config.rs config/env/CLI resolution + price table src/stats.rs per-task cost accounting (stats.tsv) src/main.rs CLI, headless mode bench/ harness, fixtures, raw data
Terminal

A terminal UI
worth reading.

Full markdown rendering in assistant output — headings, fenced code with a gutter, lists with hanging indents, blockquotes, rules and boxed pipe tables — behind turn markers, a braille activity line and a segmented status bar.

EnterSend — queued while a run is in flight
AltEnterNewline (ShiftEnter on kitty-protocol terminals)
EscInterrupt the run, or clear the input
CtrlCCtrlDQuit
CtrlLClear scrollback
CtrlAEUKWEmacs-style line editing
↑↓History
PgUpPgDnScroll · the mouse wheel works too
❯ what does handle() do in src/routes.py? ● One line: it validates the request body and dispatches to the matching route handler. ─── status ────────────────────────────────── deepseek-flash │ 60k ctx │ task $0.0002 │ $0.0011/✓ ████████▒▒▒▒▒▒▒▒▒▒ generating… ┌ rules ────────────────────────────────────┐ │ headings · lists · quotes · `inline code` │ │ tables, fenced blocks, links, strike … │ └───────────────────────────────────────────┘
Get started

One binary.
Any OpenAI-compatible model.

Point it at OpenAI, DeepSeek, OpenRouter, Ollama, vLLM, LM Studio — anything that speaks the chat-completions wire format. Configuration resolves CLI flag → environment → ~/.config/ragent/config.toml → defaults.

OpenAIDeepSeekOpenRouterOllama vLLMLM StudioAny OpenAI-compatible endpoint
# build
cargo build --release

# configure
export RAGENT_API_KEY=sk-…   # or OPENAI_API_KEY
ragent --init                # write a commented example

# run
./target/release/ragent                 # TUI
./target/release/ragent "fix the bug in src/api.rs"
git diff | ./target/release/ragent "review this"   # piped stdin
# ~/.config/ragent/config.toml
api_key = "…"
base_url = "https://api.deepseek.com/v1"
model = "deepseek-flash"
context_budget = 60000     # compaction above this
max_output_tokens = 8192
bash_timeout_ms = 30000
price_input  = 0.15        # $/1M for unknown models
price_output = 0.60
price_cached = 0.003

Linux builds link with mold by default (apt install mold); the x86_64 target is x86-64-v3. On macOS or Windows the config is a no-op.