Skip to content

Usage ledger, metering, admin API

Modulegw.07 · build · Go · Pass 7 · 4 to 6 h
You buildgo/gateway/ledger/: ledger.go (the SQLite store: Open, Record, Query), meter.go (the Meter stage of the chain), admin.go (Admin, the whole /admin/v1 surface, and UsageHandler); go/gateway/gateway.go (Deps and Admin: the stages assembled into the chain and the admin API); then the usage verb of your <system> CLI (learner territory)
Contractthe schema course/contracts/formats/usage.v1.sql (applied verbatim) and the server side of course/contracts/openapi/admin.v1.yaml; the Go API is section 4
Testscourse/tests/go/gw_07/ (what they check: section 4)
Needsgw.01 server skeleton (the chain, Exchange, WriteError), gw.02 API keys and scopes (the Principal), gw.03 rate limiting, gw.05 routing (the route stage and its admin handlers), gw.06 response cache (the stage and its purge handler), lang.11 SQL and SQLite (reading)
Used bydep.01 gateway image (pure-Go SQLite, so the image needs no C toolchain) · later ag.04 the agent’s query_usage tool, craft.14 per-version usage, ops.09 per-tenant usage
MilestoneMS-gateway (the ledger reconciles with the engine’s usage)
Optional depthSQLite: write-ahead logging (free), SQLite: atomic commit (free), SQLite: the query planner (free), OpenAI: streaming usage (free)
  • The meter is the innermost stage (proxy -> meter): it sees the finished exchange and writes one row per request, refused requests included, when the response ends.
  • An engine sends a stream’s usage only when asked: the meter sets stream_options.include_usage for every stream and hides that usage chunk from a client that never asked for it.
  • The ledger write must outlive the request: a client that hangs up mid-stream still used the tokens, and a write made on the request’s context is cancelled with it.
  • SQLite in WAL mode with one statement per row makes a SIGKILL leave every row complete or absent; busy_timeout makes concurrent writers wait instead of failing; request_id makes a retried write a no-op.
  • Queries are half-open windows [since, until) with bound parameters; a GROUP BY column comes from a whitelist, because a column name cannot be a parameter.
  • The admin API fails closed: no principal is a 401, no admin scope a 403, and a path whose module is not built yet is a well-formed 404.
  • gateway.Deps and gateway.Admin assemble the stages, so the chain’s wiring is code a test can call; the cache keys answers by the route epoch, so a new route table never serves an old answer.
Terminal window
ol start gw.07 # writes go/gateway/ledger/{ledger,meter,admin}.go and go/gateway/gateway.go with stub bodies
ol tests gw.07 # read the test catalog first
cd go && go get modernc.org/sqlite@v1.39.1 # the one third-party package (contracts/allowed-deps.toml)
ol check gw.07 # exit code is the verdict
ol diff gw.07 # after passing: your code against the reference

Then call gateway.Deps and gateway.Admin from your go/cmd/gateway composition root and add usage to your <system> CLI (section 4, “Your entry points”). MS-gateway reconciles your ledger with your engine’s usage.


After gw.01 to gw.06 your gateway authenticates every request, limits it, caches it, and routes it, and it forgets each one the moment the response ends. The rate limiter’s token buckets live in memory and reset on every restart, so they cannot say how many tokens acme used yesterday. Nothing can: not the bill, not the per-tenant panel that ops.09’s noisy-neighbor drill needs, not the agent’s query_usage tool (ag.04), and not MS-gateway, which checks that what the gateway charged equals what the engine produced. This module adds the one durable record of who used what: a row per request in a SQLite file, written by the last stage of the chain, read through GET /admin/v1/usage. It also gives the admin endpoints that gw.02, gw.05, gw.06, and gw.08 build one front door.

2.1 Metering: one row per request, from the response

Section titled “2.1 Metering: one row per request, from the response”

The chain of gw.01 is requestid -> otel -> recover -> authn -> policy -> ratelimit -> cache -> route -> proxy -> meter. The meter wraps the proxy directly, so it runs after every other stage has decided and it sees the response go by. It meters the three inference routes (/v1/chat/completions, /v1/completions, /v1/embeddings) and lets every other path through untouched. Each column of the row has one source:

ColumnSource
request_id, ts_msthe Exchange (gw.01): RequestID, Start
tenant, key_idthe Principal that gw.02’s authn stage put in the context
model, route, streamthe parsed request (Exchange.Request) and the path
served_modelthe model field of the response: an alias or a canary answers with another id
status, error_codethe status sent to the client; the error body’s code (its type when the code is null); an SSE error event’s code for a stream that failed after its first byte
prompt_tokens, completion_tokens, cached_tokensthe response’s usage object (prompt_tokens_details.cached_tokens when present)
ttft_ms, e2e_msfirst content byte minus start, NULL when none was sent or the request was refused; end minus start, on the gateway’s clock
worker_id, cache_hit, trace_idthe router’s choice and the cache flag on the Exchange; the trace id of the request’s traceparent

Where usage comes from. A non-streamed answer carries usage in its body. A stream carries it in one extra chunk near the end, {"choices": [], "usage": {...}}, and the engine sends that chunk only when the request says "stream_options": {"include_usage": true} (openai-subset.v1.yaml). Most clients do not ask, so a meter that only reads would bill every stream zero tokens. The meter therefore adds include_usage to every stream request it forwards. That changes the answer the client receives, so the meter also removes that usage chunk on the way back when the client did not ask for it. A client that did ask gets the engine’s bytes unchanged.

The invariant. For every tenant, the ledger and the engine agree:

SymbolMeaningType
tta tenantstring
LtL_tthe ledger rows of tenant ttset of rows
EtE_tthe requests of tt that the engine answered with a usage objectset of requests
pr,crp_r, c_rprompt and completion tokens of a row or request rrintegers ≥0\ge 0

∑r∈Ltpr=∑e∈Etpe,∑r∈Ltcr=∑e∈Etce\sum_{r \in L_t} p_r = \sum_{e \in E_t} p_e, \qquad \sum_{r \in L_t} c_r = \sum_{e \in E_t} c_e

Refused requests are rows with zero tokens, so they count in requests and errors and add nothing to the sums. TestLedgerSumEqualsEngineUsage checks the invariant over 1000 requests, and MS-gateway checks it against your real engine.

2.2 The store: SQLite, WAL, one statement per row

Section titled “2.2 The store: SQLite, WAL, one statement per row”

The ledger is a single SQLite file at [gateway].usage_db. Open applies usage.v1.sql exactly as written (every statement is IF NOT EXISTS or OR IGNORE, so reopening is a no-op), and refuses a file whose schema_version is newer than 1: an old gateway must not write into a table a later migration reshaped. The driver is modernc.org/sqlite, SQLite compiled to pure Go, so the gateway still builds with CGO_ENABLED=0 into a static binary (dep.01).

Four settings decide whether rows survive (lang.11 worked through each one):

SettingScopeWhy
journal_mode = WALstored in the filea commit appends frames and a commit record to usage.db-wal; readers never block the writer; recovery ignores a tail without its commit record
synchronous = NORMALper connectionin WAL mode, a commit survives a process kill (the bytes are in the OS page cache); only a power loss can take the last commits. FULL would fsync on every row
busy_timeout = 5000per connectionSQLite has one writer at a time; without a timeout a second writer fails at once with SQLITE_BUSY and its row is lost
request_id primary key, ON CONFLICT DO NOTHINGschema and statementa retried write (a timeout after the commit) leaves the first row alone

database/sql keeps a pool of connections, so a per-connection setting made with one Exec("PRAGMA ...") reaches one connection only. The settings go into the data source name instead, which the driver applies to every connection it opens: usage.db?_pragma=busy_timeout(5000)&_pragma=journal_mode(WAL)&_pragma=synchronous(NORMAL).

A row is one INSERT statement, so it commits whole or not at all. Record returns only after the commit: an acknowledgement means the row is in the WAL. Cached tokens larger than the prompt (an engine bug) would violate the table’s CHECK and lose the whole row, so Record clamps them; negative counts and an empty request id are refused with ErrInvalidRecord.

2.3 Queries: half-open windows, bound values, whitelisted columns

Section titled “2.3 Queries: half-open windows, bound values, whitelisted columns”

Query(UsageFilter) sums rows: requests, errors (status at least 400, or an error code), and the three token counts.

  • Windows are half-open, since <= ts_ms < until. Then [10:00, 11:00) and [11:00, 12:00) add up to [10:00, 12:00) with no request at exactly 11:00:00.000 counted twice.
  • Values are bound parameters (tenant = ?). The tenant arrives from the admin API’s query string; spliced into the SQL text, x' OR '1'='1 reads every tenant’s usage.
  • A column name cannot be a parameter, so group_by is looked up in a fixed map (model, tenant, key_id, api_version) and anything else is ErrBadGroupBy.
  • No grouping gives exactly one row, zeros when nothing matches, so a caller never has to special-case an empty answer. Grouped rows are sorted by the group key.

The kubelet SIGKILLs a pod that exceeds its memory limit or fails its liveness probe; nothing in the process runs afterwards. What survives is whatever reached the OS before the signal. With WAL and one statement per row, every committed row is in usage.db-wal (the OS writes it to disk later even though the process is gone), and a commit that was half-written has no commit record, so the next Open ignores it. Two designs break this: a ledger that collects rows in memory and writes them in batches (every acknowledged row in the batch is lost), and one that writes a row in two statements without a transaction (a kill between them leaves a torn row). TestSIGKILLLeavesRowsCompleteOrAbsent kills a writer process after 400 acknowledgements and checks every acknowledged row, column by column.

2.5 The admin API: one front door, fail closed

Section titled “2.5 The admin API: one front door, fail closed”

admin.v1.yaml has eight operations owned by five modules. This module owns the server side: one handler, Admin(AdminDeps), for every /admin/v1 path. It serves usage itself and hands every other path to the handler its module builds (keys from gw.02, workers, routes, and drains from gw.05, the cache purge from gw.06, the policy from gw.08). A path whose handler is nil answers 404 not_found, so a gateway in the middle of Pass 7 still has a well-formed admin surface. Admin checks the principal itself: no principal is 401 invalid_api_key and a key without admin is 403 insufficient_scope. gw.02’s authn stage already enforces that, but an admin handler that trusts its mount point serves every tenant’s billing data the first time someone mounts it outside the chain.

Four requests reach the meter (times are 2026-10-01, UTC):

request_idstarttenantkeymodelroutestatuscodepromptcompletion
req-112:00acmeacmekeyaaaaasmolchat, stream2001230
req-212:01acmeacmekeybbbbbsmolchat2002010
req-312:02acmeacmekeyaaaaatinycompletions429rate_limit_exceeded00
req-412:03globexglobexkeycccsmolchat20075

The stream, req-1. The client sends {"model":"smol","stream":true,"messages":[{"role":"user","content":"Once"}]}. The meter forwards {"messages":[...],"model":"smol","stream":true,"stream_options":{"include_usage":true}}. The engine answers with three content chunks, then data: {...,"choices":[],"usage":{"prompt_tokens":12,"completion_tokens":30,"total_tokens":42}}, then data: [DONE]. The meter reads 12 and 30 from the usage chunk, drops that chunk, and the client receives the three content chunks and [DONE]: exactly what it would have received from an engine it asked directly.

The rows, then the queries.

  1. acme, no grouping: the rows are req-1, req-2, req-3. Requests 33; errors 11 (req-3, status 429); prompt 12+20+0=3212 + 20 + 0 = 32; completion 30+10+0=4030 + 10 + 0 = 40.
  2. acme, grouped by model, sorted by the key: smol is req-1 and req-2, so requests 2, errors 0, prompt 32, completion 40; tiny is req-3, so requests 1, errors 1, tokens 0.
  3. The window since=12:01, until=12:03, every tenant: 12:01 <= ts < 12:03 holds for req-2 and req-3 and not for req-4 (12:03 is not before 12:03). Requests 2, errors 1, prompt 20, completion 10.
  4. initech: no rows, so one row of zeros.

Through the admin API, query 3 grouped by model is

GET /admin/v1/usage?since=2026-10-01T12:01:00Z&until=2026-10-01T12:03:00Z&group_by=model
Authorization: Bearer tl_<admin key>
200 {"since":"2026-10-01T12:01:00Z","until":"2026-10-01T12:03:00Z","data":[
{"model":"smol","requests":1,"errors":0,"prompt_tokens":20,"completion_tokens":10,"cached_tokens":0},
{"model":"tiny","requests":1,"errors":1,"prompt_tokens":0,"completion_tokens":0,"cached_tokens":0}]}

These numbers are TestHandWorkedExample, TestStreamUsageIsRequestedAndHidden, and TestAdminUsageEndpoint.

package ledger // import "tinyllm/gateway/ledger"
const Schema = `...` // contracts/formats/usage.v1.sql, byte for byte (written by ol start)
const SchemaVersion = 1
type UsageRecord struct {
RequestID string; Start time.Time; Tenant, KeyID, Model, ServedModel, Route, APIVersion string
Status int; ErrorCode string; Stream bool
PromptTokens, CompletionTokens, CachedTokens int
TTFT *time.Duration; E2E time.Duration; CacheHit bool; WorkerID, TraceID string
}
type UsageFilter struct { Tenant, KeyID string; Since, Until time.Time; GroupBy string }
type UsageRow struct { Tenant, KeyID, Model, APIVersion string; Requests, Errors, PromptTokens, CompletionTokens, CachedTokens int64 }
type Ledger interface {
Record(ctx context.Context, r UsageRecord) error
Query(ctx context.Context, f UsageFilter) ([]UsageRow, error)
}
var ErrInvalidRecord, ErrBadGroupBy, ErrNewerSchema error
func DSN(path string) string
func Open(path string) (*DB, error) // *DB implements Ledger; Close() error
func WithIncludeUsage(body []byte) ([]byte, bool, error)
func Meter(l Ledger, o MeterOptions) server.Middleware // MeterOptions{Clock, WriteTimeout, Logger}
func Admin(d AdminDeps) http.Handler // AdminDeps{Usage Ledger; Keys, Workers, Routes, Drain, Policy, Cache http.Handler}
func UsageHandler(l Ledger) http.Handler // GET /admin/v1/usage
package gateway // import "tinyllm/gateway"
type Stages struct {
Keys *auth.Store; Logger *slog.Logger // gw.02
Policy server.Middleware; PolicyAdmin http.Handler // gw.08
Limiter *limit.Limiter; Counter limit.Counter // gw.03
Cache *cache.LRU; CacheOptions cache.Options // gw.06
Registry *route.Registry; Router *route.Router // gw.05
Proxy http.Handler // route.Proxy, or gw.04's proxy
Ledger ledger.Ledger; Meter ledger.MeterOptions // this module
}
func Deps(s Stages) server.Deps // a nil stage is left out; the cache's Rev defaults to route.ETag(epoch)
func Admin(s Stages) ledger.AdminDeps // KeysHandler, WorkersHandler, RoutesHandler, DrainHandler, PurgeHandler

gateway.go is the part of the composition root a test can call. Deps calls each module’s middleware constructor (auth.Middleware, limit.Middleware, cache.Middleware, route.Middleware, ledger.Meter) and leaves a nil stage out; gw.01’s Handler puts them in contract order. One wiring decision is made here and nowhere else: with a Router and no CacheOptions.Rev, the cache revision is the route table’s ETag, so a PUT /admin/v1/routes makes every answer cached under the old table unreachable. Admin builds each admin path’s handler from the module that owns it.

ol start gw.07 writes the four files with every function body stubbed (panic("todo: gw.07")); types, constants, and Schema stay. The unexported helpers (meterWriter, inspect, errorCode, traceID, …) are a suggested decomposition.

Your entry points. In go/cmd/gateway, open the ledger at [gateway].usage_db, build the stages into a gateway.Stages, and serve the admin API through a second, short chain that has only authn in front of it, so admin calls never reach the router:

st := gateway.Stages{Keys: keys, Limiter: limiter, Counter: counter, Cache: lru, CacheOptions: cache.Options{TTL: ttl},
Registry: reg, Router: router, Proxy: route.Proxy(router, route.Options{}), Ledger: db}
inference := server.New(cfg, gateway.Deps(st)).Handler()
admin := server.New(cfg, server.Deps{Keys: auth.Middleware(keys, nil), Proxy: ledger.Admin(gateway.Admin(st))}).Handler()
mux := http.NewServeMux()
mux.Handle("/admin/v1/", admin)
mux.Handle("/", inference)

and add usage to your <system> CLI: <system> usage --tenant acme --since 1h [--group-by model] calls GET /admin/v1/usage with the admin key from TL_API_KEY and prints the rows.

TestKINDChecksWhy it matters downstream
TestHandWorkedExampleunitsection 3: acme 3/1/32/40, by model, the window, the empty tenantyou and the tests agree on every definition
TestSchemaMatchesContractconformanceOpen creates what usage.v1.sql creates (compared through sqlite_master), version 1, WALag.04 and ops.09 read these tables
TestReopenKeepsRowsAndRefusesNewerSchemaboundaryreopening keeps rows; schema_version 2 is refused with ErrNewerSchemarollouts restart the gateway; craft.14 migrates the schema
TestRecordIsIdempotentByRequestIDboundarya second record of req-1 changes nothingretried writes never double bill
TestRecordValidatesAndClampsCachedTokensboundarycached 50 > prompt 20 is stored as 20; negative counts and an empty id are refusedan engine bug never loses a row
TestRecordStoresEveryColumnunitevery field in its column and unit; NULL code and TTFT; api_version defaults to '1'columns, not Go structs, are the contract
TestQueryWindowIsHalfOpenboundary[10:00,11:00) + [11:00,12:00) = [10:00,12:00)hourly panels add up
TestQueryValuesAreBoundParametersunitquotes in tenant or key id change nothingthe admin API passes them from the URL
TestQueryGroupByIsWhitelistedboundarystatus, Model, an injection are ErrBadGroupBy; key rows sortedthe admin API’s 400
TestConcurrentRecordsAllLandfault16 writers x 64 rows, no SQLITE_BUSY, 1024 rowsevery request goroutine records at once
TestHelperLedgerWriterfaultnot a check: the child process of the next test (skips otherwise)
TestSIGKILLLeavesRowsCompleteOrAbsentfaultafter SIGKILL: integrity ok, WAL, every acknowledged row present and completeOOM kills on kind; drill ops.01
TestLedgerSumEqualsEngineUsageproperty1000 mixed requests: per-tenant ledger sums equal the engine’s usage; 59 refusals countedMS-gateway’s reconciliation
TestStreamUsageIsRequestedAndHiddenunitinclude_usage added and its chunk hidden; a client that asked gets the bytes unchangedstreams are billed; clients see what they asked for
TestMeterDoesNotBufferTheStreamunitthe first event reaches the client while the engine holds the restTTFT, the gateway’s SLO
TestRefusedRequestsAreRecordedWithTheirCodeunit429 and 400 rows with code, zero tokens, NULL TTFT; a mid-stream error event is an errorops.09 sees who hits limits
TestMeterRecordsWhoWhatAndWhereunittenant, key, model, served model, route, worker, cache flag, trace id, gateway-clock latencyper-tenant and per-worker panels
TestClientDisconnectStillRecordsfaulta client that hangs up mid-stream still gets a rowabandoned streams still cost tokens
TestUnmeteredPathsPassThroughboundaryGET /v1/models untouched, no rowonly inference is billed
TestAdminUsageEndpointconformance{since, until, data}, null window, one ungrouped row, rows grouped by model<system> usage, ag.04
TestAdminUsageRejectsBadParametersboundarybad since, until, group_by, an empty window: 400 naming the param; POST is 405the CLI can name the bad flag
TestAdminRequiresAdminScopeunitno principal 401, an infer key 403billing data stays private
TestAdminMountsEveryPathuniteach path reaches its module’s handler; nil handlers and unknown paths are 404 not_foundgw.02, gw.05, gw.06, gw.08 mount here
TestGatewayAssemblesTheStagesunitreal keys, limits, cache, router, and ledger through gateway.Deps and gateway.Admin: a repeat is a hit, a route PUT makes it a miss, the 4th request of an RPM-3 key is a 429 before the cache, purge, workers, routes, drain, keys, and usage reach their handlersyour go/cmd/gateway is these two calls
PitfallSymptomCaught by
1. reading usage only from the responseevery stream is billed 0 tokens: engines send stream usage only when askedTestStreamUsageIsRequestedAndHidden, TestLedgerSumEqualsEngineUsage (mutant s02)
2. writing the row on r.Context()a client that hangs up mid-stream leaves no row: the write is cancelled with the requestTestClientDisconnectStillRecords (mutant s01)
3. no busy_timeout, or a PRAGMA run once on a pooled *sql.DBdatabase is locked under load, rows lostTestConcurrentRecordsAllLand (mutant s06)
4. INSERT OR REPLACE (or an upsert) on request_ida retried write rewrites the first row with new numbersTestRecordIsIdempotentByRequestID (mutant s04)
5. wrapping the ResponseWriter without a working Flushtokens wait in a 4 KiB buffer; TTFT becomes the whole answer’s timeTestMeterDoesNotBufferTheStream (mutant s20)
6. trusting the engine’s cached_tokensthe CHECK (cached_tokens <= prompt_tokens) rejects the row and it is lostTestRecordValidatesAndClampsCachedTokens (mutant s10)
7. building SQL with fmt.Sprintftenant=x' OR '1'='1 reads every tenantTestQueryValuesAreBoundParameters (mutant s08)
8. an admin handler that trusts its mount pointmounted outside authn, it serves billing data to anyoneTestAdminRequiresAdminScope (mutant s13)
9. acknowledging before the commit, or no WALa SIGKILL loses acknowledged rows, or a rollback journal blocks readersTestSIGKILLLeavesRowsCompleteOrAbsent, TestSchemaMatchesContract (mutants s16, s05)
10. forwarding the usage chunk you asked forclients that never asked for usage receive a chunk with choices: []TestStreamUsageIsRequestedAndHidden (mutant s03)
11. until inclusivethe request at exactly 11:00 is in two hourly windowsTestQueryWindowIsHalfOpen (mutant s07)
12. recording only successes, or a TTFT for refusalsrefused traffic is invisible; error rows show a latency they never hadTestRefusedRequestsAreRecordedWithTheirCode (mutants s15, s12, s22)
13. a cache revision that ignores the route tableafter a PUT /admin/v1/routes the cache keeps serving answers from the old route’s modelTestGatewayAssemblesTheStages (mutant s23)
DirectionModuleHow it uses this
Backgw.01the chain the meter closes, the Exchange it reads, WriteError for the admin errors
Backgw.02the Principal: tenant and key id per row, the admin scope Admin checks
Backgw.03gateway.Deps puts limit.Middleware in the Limiter slot
Backgw.05gateway.Deps puts route.Middleware in the chain and keys the cache by route.ETag; gateway.Admin builds the workers, routes, and drain handlers
Backgw.06gateway.Deps puts cache.Middleware in the chain; gateway.Admin builds PurgeHandler
Backlang.11the SQL: transactions, WAL, indexes, aggregates, and the same usage table
Forwarddep.01builds the gateway image with CGO_ENABLED=0 around the pure-Go driver, the usage DB on a volume
ForwardMS-gatewayreconciles your ledger against your engine’s usage through the admin API
Forwardag.04 (B12), craft.14, ops.09the agent’s query_usage tool reads these tables; API v2 fills api_version and cached_tokens; the noisy-neighbor drill reads per-tenant usage

If you skip this module, MS-gateway’s reconciliation step fails and ol check dep.01 reports needs gw.07.

Your pieceProduction equivalentWhat it addsWhere to look
the SQLite ledgerLitestream, ClickHouse, OpenMeterstreaming the WAL to object storage; columnar usage analytics; metering as an event pipelineLitestream (free), OpenMeter (free)
MeterLiteLLM spend tracking, Envoy AI Gateway token usageper-key budgets enforced from the same usage, cost per modelLiteLLM spend tracking (free), Envoy AI Gateway usage limiting (free)
the admin APIStripe usage records, Kong Admin APIidempotency keys on writes, pagination, audit logsStripe idempotent requests (free)