Skip to content

Usage policy at the gateway

Moduleethics.05 · practice · policy and documentation · Pass 10 · 1 to 2 h
You builddocs/USAGE_POLICY.md and deploy/policy.v1.yaml
Contractpolicy.v1.schema.json, linear-head.schema.json
Testscourse/tests/ethics.05/ (why: policy ordering, classifier decisions, red-team coverage, and audit privacy)
Needsgw.08 usage policy enforcement, L6.5 linear head, data.05 privacy detectors
Used byMS-agent checks the policy matrix and red-team prompts
MilestoneMS-agent
Optional depthNIST AI Risk Management Framework, Govern and Measure functions
  • Policy rules are ordered and the first matching rule wins.
  • Model and tenant restrictions are explicit; classifier rules use the engine embedding endpoint and a versioned linear head.
  • A classifier or embeddings failure fails closed with 503.
  • Every deny has a stable rule id, a useful reason, and a privacy-safe audit record.
  • The policy describes permitted use; it does not claim the classifier can determine intent or replace human review.
Terminal window
ol start ethics.05
ol tests ethics.05
ol check ethics.05

Your agent now reaches users through the gateway, and its tool gates cannot prevent every harmful inference request. A policy layer gives operators a place to set model, tenant, token, and classifier restrictions, and gives users a consistent explanation when a request is blocked.

The policy document is configuration, not a promise that a model is safe. Keep rules narrow, explain their scope, and review false positives and false negatives. The linear head scores normalized embeddings using z = W e + b; the softmax probability for class c is exp(z_c) / sum_j exp(z_j). Its training labels, threshold, embedding model, and held-out metrics belong in the model card. The Go gateway evaluates the exported head so Python is not in the request path.

SymbolMeaning
enormalized embedding returned by engine /v1/embeddings
Wclassifier weights, one row per class
bone bias per class
zclass logits
τprobability threshold that activates a classifier rule

Suppose a request has max_tokens = 2048, tenant free, and matches a rule capped at 1024 tokens. That rule denies before any classifier call. For a separate request, logits [0, 2] give unsafe probability e²/(1+e²) ≈ 0.881; with threshold 0.9, that classifier rule allows it. The fixtures exercise both cases, including the embeddings service failing.

Write USAGE_POLICY.md with scope, responsible owner, appeal path, review cadence, model and data limitations, and the actions users can expect. Define ordered rules in the schema. Include model restrictions, a token cap, and the fitted classifier head at models/smol-135m/heads/usage.json. Explain why each red-team fixture is denied without copying sensitive user content into audit logs.

TestWhy it existsExpected result
test_policy_schema_and_orderRejects ambiguous or misspelled rulesFirst matching deny wins
test_red_team_prompts_are_deniedChecks configured refusals at the policy boundaryEvery blocked fixture returns 451 with its rule id
test_no_pii_in_audit_examplesAvoids retaining prompt PII in logsRule id and request id only; redact detector matches
PitfallCaught by
Putting a broad allow rule before a targeted denytest_policy_schema_and_order; mutant s01
Treating an embeddings timeout as safetest_red_team_prompts_are_denied; mutant s02
Logging the prompt in a deny eventtest_no_pii_in_audit_examples; mutant s03
Applying the threshold to a raw logittest_policy_schema_and_order; mutant s04
DirectionModuleHow it uses this
Backgw.08Read the enforcement behavior this policy artifact configures.
BackL6.5Inspect how the exported linear head is fitted and represented.
Backdata.05Reuse its privacy detector categories when writing safe audit examples.
ForwardMS-agentChecks the configured policy matrix, red-team fixture set, audit log, and gateway behavior.
Forwardfield.06Uses this policy in a customer security review.

Production policy systems add policy version rollout, appeals, jurisdiction-specific rules, monitoring for drift, and human review for high-impact decisions. Keep each addition traceable to an owner and an observed failure mode.