Skip to content

Usage policy enforcement

Modulegw.08 · build · Go · Pass 10 · 2 to 3 h
You buildgo/gateway/policy/policy.go
Contractpolicy.v1.schema.json, linear-head.schema.json
Testscourse/tests/go/gw_08/ (why: ordered rules, linear-head parity, audit, and fail-closed behavior)
Needsgw.01 middleware chain, gw.02 principal, gw.05 route.InferenceRequest, ethics.05 policy, data.05 detector list
Used bygw.07 route layer consumes policy decisions; MS-agent exercises the policy fixtures
MilestoneMS-agent
Optional depthOPA and Cedar policy evaluation
  • Evaluate policy rules in file order and use the first match.
  • Model, tenant, and token limits are cheap checks; classifier rules use engine embeddings.
  • Stable rule ids travel through audit events and the 451 response.
  • Classifier dependencies fail closed with 503.
  • Redact request text with the Go PII detector port before logging.
Terminal window
ol start gw.08
ol tests gw.08
ol check gw.08

The agent now sends user requests through the gateway, so a usage restriction must be enforced on every inference path. Middleware placement after authentication gives the policy a tenant and model identity while keeping it ahead of rate limiting, cache, routing, and proxying.

Policy.Check(ctx, principal, request) returns an allow or deny decision. A rule matches only when all stated match fields match. max_tokens: N matches a request whose requested output cap exceeds N. A classifier rule calls the configured engine’s /v1/embeddings, normalizes the returned vector, evaluates the exported head, and compares the selected class probability with its threshold. A matching deny is audit-logged and returned as 451 usage_policy; dependency failure is 503.

For an unsafe class logit of 2 and a safe logit of 0, p(unsafe) = exp(2)/(exp(0)+exp(2)) ≈ 0.881. The decision is allow at threshold 0.9 and deny at threshold 0.85. TestPolicyOrderingAndModelTenantRules also walks a hand-built request through ordered rules. A zero embedding produces the softmax of the bias vector, not NaN. The large-vector fixture ensures normalization and dot products remain finite.

Implement the middleware consumed by server.Deps.Policy. Use the exchange’s parsed request and authenticated principal. Do not log the request body. The named checks are TestPolicyOrderingAndModelTenantRules, TestLinearHeadFixtureParity, TestZeroAndLargeEmbeddingsAreFinite, TestEmbeddingsFailureFailsClosed, TestEveryDenyIsAudited, and TestRedactPII.

TestWhy it existsExpected result
TestPolicyOrderingAndModelTenantRulesProtects rule semanticsStable first matching decision
TestPolicyYAMLLoadChecks the versioned operator policy artifactLoads valid policy and rejects malformed YAML
TestLinearHeadFixtureParityProtects Python/Go boundaryDifference at most 1e-6
TestZeroAndLargeEmbeddingsAreFiniteKeeps normalization finite at boundary valuesFinite scores for zero and large vectors
TestEmbeddingsFailureFailsClosedAvoids fail-open behavior503 and audit event
TestEveryDenyIsAuditedMakes both rule and default denials durableOne audit event per deny with the matching rule id
TestRedactPIIProtects audit privacyRedacted prompt data
PitfallCaught by
Matching max_tokens in the opposite directionTestPolicyOrderingAndModelTenantRules; mutant s01
Reordering rulesTestPolicyOrderingAndModelTenantRules; mutant s02
Using raw logits as probabilitiesTestLinearHeadFixtureParity; mutant s03
Allowing an embeddings errorTestEmbeddingsFailureFailsClosed; mutant s04
Returning a deny without recording its rule idTestEveryDenyIsAudited
Auditing raw prompt textTestRedactPII; mutant s05
DirectionModuleHow it uses this
Backgw.01Places the policy middleware in the authenticated request chain.
Backgw.02Supplies the principal used for tenant and model matching.
Backgw.05route.InferenceRequest is the parsed request the policy matches on.
Forwardgw.07Consumes policy decisions in the route layer.
ForwardMS-agentChecks usage restrictions through the public gateway interface.
Your pieceProduction equivalentWhat it addsWhere to look
ordered rulessigned policy revisions and rollout controlsAuditable change history and staged activationOPA and Cedar policy evaluation
451 decisionmeasured false-positive reviewOperator feedback and correction loopsEnvoy AI Gateway policy examples