Adapted from the chitin-api repository docs. Private repo.
Hosted AI gateway that optimizes every token — caching, semantic dedup, context compression, and intelligent model routing. Change one URL in your code and get ≥40% token savings across all LLM calls. (Related: chitin product repo, chitin-dashboard.)
Your App (OpenAI SDK)
│ base_url = "https://api.usechitin.com/v1"
▼
┌──────────────────────────────────┐
│ Chitin │
│ Auth → Budget check → Cache │
│ → Semantic dedup → Compression │
│ → Model routing → Dispatch │
└──────────────────┬───────────────┘
│ (using YOUR provider API keys)
┌──────────┼──────────┐
▼ ▼ ▼
Anthropic OpenAI Any OAI-compatible
cache_control)| Plan | Monthly | Proxy Fee | Included Tokens |
|---|---|---|---|
| Free | $0 | — | 1M tokens/mo |
| Pro | $49 | $0.20/MTok | 50M tokens base |
| Scale | $249 | $0.15/MTok | 300M tokens base |
| Enterprise | Custom | Custom | Unlimited |
Proxy fee applies only to tokens actually sent upstream — saved tokens are never charged.
X-Gateway-Session-Id (scope budget/cache) · X-Gateway-Tag (routing/batching hints) · X-Gateway-No-Cache · X-Gateway-No-Route · X-Gateway-Compact (manual compaction)
Every response includes transparency headers: X-Gateway-Tokens-Saved, X-Gateway-Cache (hit/miss/dedup), X-Gateway-Model (actually used), X-Gateway-Model-Routed, X-Gateway-Compression, X-Gateway-Cost-USD.
Go 1.23+ · PostgreSQL 16 + pgvector · Redis 7 · Kubernetes (stateless, autoscaled) · MiniLM-L6 via ONNX in-process · Stripe Meters API for billing.
black-candle-technologies/chitin-api (private)https://api.usechitin.com/v1 · App: https://app.usechitin.com