Talanton (τάλαντον)

Open-source AI cost tracking, pre-flight token metering & LLM budget guardrails

v0.2.0 on PyPIpythonsub-2ms guardrailslocal-first60+ modelsPyPI (talanton-py) →Live Docs →Repository →
Talanton — Classical Engraving of Icarus: Preventing Runaway AI Overspend
FIG. 01 // WHY TALANTON EXISTS — STOPPING ACCIDENTAL $1,000+ AI BILLSSUB-2MS GUARDRAIL

The Core Problem

When building modern AI-driven applications, prompt costs silently compound. An unmonitored while-loop in an autonomous agent, an unexpected recursion bug, or an employee accidentally submitting a 100-page PDF into an expensive model can burn through thousands of dollars over a single weekend.

Existing monitoring tools (such as Helicone or Langfuse) require proxying all requests through third-party remote clouds. This introduces 40ms to 120ms of artificial network latency to every prompt and forces proprietary user data across the open internet.

The Solution: Local-First Budget Guardrails

Talanton (τάλαντον) is named after the ancient Greek balance scale of honest weight. It runs 100% in-process on the local machine with sub-2ms latency and zero cloud data exfiltration.

Before any request is sent to OpenAI or Anthropic, Talanton meters the exact token count, calculates cost in real time across 60+ frontier models, and compares the projected spend against pre-configured soft warnings and hard circuit-breaker limits.

from talanton import TalantonTracker, BudgetGuardrail from talanton.guardrails import raise_exception tracker = TalantonTracker() # Stop calls if the daily budget hits $10.00 guardrail = BudgetGuardrail( tracker=tracker, soft_limit=5.00, hard_limit=10.00, period="day", on_hard_limit=raise_exception ) # Checks spend in 0.08ms before any outbound API call is dispatched result = guardrail.check(estimated_cost=0.04) if result.is_blocked: print(f"Call blocked! Daily spend reached ${result.current_spend:.2f}")

Drop-in 1-Line SDK Wrappers

Developers can wrap existing standard OpenAI and Anthropic clients with a single line of code. Every chat completion is metered, recorded to an embedded SQLite WAL ledger, and enforced against spending limits with zero boilerplate:

import openai from talanton import TalantonTracker, BudgetGuardrail from talanton.integrations import TalantonOpenAI tracker = TalantonTracker() guard = BudgetGuardrail(tracker=tracker, hard_limit=25.0, period="day") # 1-Line Drop-in Client Replacement: client = TalantonOpenAI(openai.OpenAI(), tracker=tracker, guardrail=guard) # Standard OpenAI syntax — automatically guarded & tracked: response = client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "Analyze quarterly filings"}] )

Performance & Scalability Benchmarks

Measured under 25-thread concurrency on Python 3.12 with SQLite WAL persistence:

  • Pre-Check Latency (P50): 0.08 ms (1,000x faster than cloud proxies).
  • Pre-Check Latency (P99): 1.94 ms (strict sub-2ms SLA).
  • Ingestion Throughput: 24,586 calls/sec with zero lock contention.
  • Data Privacy: 100% local-first — customer prompts never leave your infrastructure.
  • Model Registry: Exact pricing and tokenization for 61 models (GPT-4o, Claude 3.7 Sonnet, Llama 3.3).