Talanton (τάλαντον)
Open-source AI cost tracking, pre-flight token metering & LLM budget guardrails
The Core Problem
When building modern AI-driven applications, prompt costs silently compound. An unmonitored while-loop in an autonomous agent, an unexpected recursion bug, or an employee accidentally submitting a 100-page PDF into an expensive model can burn through thousands of dollars over a single weekend.
Existing monitoring tools (such as Helicone or Langfuse) require proxying all requests through third-party remote clouds. This introduces 40ms to 120ms of artificial network latency to every prompt and forces proprietary user data across the open internet.
The Solution: Local-First Budget Guardrails
Talanton (τάλαντον) is named after the ancient Greek balance scale of honest weight. It runs 100% in-process on the local machine with sub-2ms latency and zero cloud data exfiltration.
Before any request is sent to OpenAI or Anthropic, Talanton meters the exact token count, calculates cost in real time across 60+ frontier models, and compares the projected spend against pre-configured soft warnings and hard circuit-breaker limits.
Drop-in 1-Line SDK Wrappers
Developers can wrap existing standard OpenAI and Anthropic clients with a single line of code. Every chat completion is metered, recorded to an embedded SQLite WAL ledger, and enforced against spending limits with zero boilerplate:
Performance & Scalability Benchmarks
Measured under 25-thread concurrency on Python 3.12 with SQLite WAL persistence:
- Pre-Check Latency (P50):
0.08 ms(1,000x faster than cloud proxies). - Pre-Check Latency (P99):
1.94 ms(strict sub-2ms SLA). - Ingestion Throughput:
24,586 calls/secwith zero lock contention. - Data Privacy: 100% local-first — customer prompts never leave your infrastructure.
- Model Registry: Exact pricing and tokenization for 61 models (GPT-4o, Claude 3.7 Sonnet, Llama 3.3).
