Skip to content

Tool calls and constrained JSON decoding in the engine

ModuleL10.9 · build · Rust · Pass 7 · 12 to 16 h
You buildrust/crates/tl-engine/src/constrain.rs: the regex subset to a canonical byte DFA (Thompson NFA, subset construction over byte classes, dead states removed, Moore minimization, breadth-first numbering), the JSON-schema subset to a regex, token masks over a byte trie, constrained_sample, an order-keeping JSON reader that writes Python’s compact JSON · rust/crates/tl-serve/src/tools.rs: tools, tool_choice, response_format validated, tools rendered into the prompt, the grammar per tool_choice, a streaming tool-call parser, and the OpenAI tool_calls message and deltas · rust/crates/tl-serve/src/server.rs, taken over from L10.7: the chat route reads the tool fields, decodes a request with a grammar under its Constraint, and answers with tool_calls
Contractthe Python specification you port: py/tinyllm/infer/constrain.pyi (L8.7) · HTTP: openapi/openai-subset.v1.yaml (tools, tool_choice, response_format, tool_calls deltas, finish_reason: tool_calls) · conformance tools.call, tools.stream, tools.choice
Testscourse/tests/rust/l10_9.rs, 20 tests (what they check: section 4); the L10.5 and L10.7 tests run as the regression of the server you take over
NeedsL8.7 your Python constrained decoding, the specification (chapter) · L10.1 the sampler (chapter) · L10.5 the chat template and the API error shape (chapter) · L1.5 the Tokenizer trait (chapter) · L10.7 the serve loop you take over, which records metrics and spans (chapter) · reading: M06.1 graphs and reachability (chapter) · or --ref-deps
Used bythe serve loop you take over, and so its call site obs.02, which runs your engine; ag.01 and MS-agent read tool_calls from it over HTTP; craft.14 takes the serve loop over next to serve API v1 and v2 side by side
MilestoneMS-L10 (and MS-agent, where SmolLM2-135M-Instruct calls your agent’s tools through it)
Optional depthHopcroft, Motwani, Ullman, Introduction to Automata Theory, ch. 2 to 4; Willard and Louf, Efficient Guided Generation (free); XGrammar (free); OpenAI function calling guide (free)
  • A minimized DFA numbered breadth-first over bytes in order is canonical: your Rust tables equal your Python’s for every pattern and schema, so the masks are the same (hand_example_dfa, dfa_matches_python_l8_7, masks_match_python_l8_7).
  • A token is allowed when walking ALL its bytes stays live; EOS only where the output is a whole string; specials never (token_index_rules).
  • Whatever the logits, constrained output parses and satisfies its schema: that is why a random tiny model can still make a valid forced tool call (constrained_output_always_parses, forced_call_from_a_random_model_is_valid).
  • tool_choice decides the grammar: a named function is one call of it, required one to four calls, auto free text until the model writes <tool_call>, none free text (grammar_by_tool_choice).
  • The parser streams argument fragments as they arrive and never leaks markup into the content, wherever the tokens split the text (parser_hand_example, parser_is_split_invariant).
Terminal window
ol start L10.9 # stubs constrain.rs and tools.rs
ol tests L10.9
ol check L10.9 # exit code is the verdict
ol check L10.9 --ref-deps # only if L8.7, L10.1, L10.5, or L1.5 is not passing yet
ol diff L10.9
ol conform openapi:v1 --target engine # tools.* turn from pending to checked once L10.9 passes

Declare pub mod constrain; in tl-engine/src/lib.rs and pub mod tools; in tl-serve/src/lib.rs. This module takes over server.rs from L10.7 (upgrades), so ol check L10.9 also runs the L10.5 and L10.7 tests as regressions; one of L10.5’s error cases is response_format json_schema, which stays 422 here. Your chat handler calls parse_tool_request on the raw body (before the generic parser, so property order survives), renders with render_prompt, builds one TokenIndex per (grammar, model) and caches it, runs each sequence through a Constraint (from the first token, or from the trigger in auto mode), and turns the text through ToolParser into delta_json chunks or message_json. A request with a grammar runs on the engine thread outside the batch, as the speculative path does: its own KV blocks through spec::RunnerTarget, constrained_sample per token, and the request’s own generator; a grammar that is complete with no token left to allow (a vocabulary without EOS) ends the answer with stop.


Agents (ag.01 onward) need structured output from your engine: a tool name and JSON arguments that a program can execute. Asking nicely in the prompt works for large models some of the time and for a 135M model almost never. In Python (L8.7) you built the machinery that makes it work every time: compile the allowed language to a DFA over bytes and mask, at every step, every token that would leave it. Your Rust engine answers tools and response_format with 422 unsupported_parameter today. This module ports the automaton to the engine (the masks must equal the Python ones exactly, so the port is held to them) and adds the OpenAI tool-calling surface around it.

SymbolMeaningType
Σ\Sigmathe 256 byte valuesalphabet
δ(s,b)\delta(s, b)the state after byte bb from state ss; ⊥\bot (DEAD) when no accepting state is reachablei32
AAthe accepting statesbool[n]
wtw_tthe bytes of token ttVec<u8>
δ∗(s,w)\delta^*(s, w)δ\delta applied byte by byte; ⊥\bot stays ⊥\boti32

The pipeline is L8.7’s, step for step. Parse the subset (literals, escapes, classes \d \w \s and negations, ., brackets, groups, |, * + ? {m} {m,} {m,n}) into a tree of byte sets. Thompson’s construction turns each node into a small NFA fragment glued with epsilon edges. The subset construction runs over byte classes (bytes every edge treats alike) instead of 256 bytes, so a JSON string pattern with its UTF-8 ranges stays small. Remove the states from which no accepting state is reachable (their transitions become ⊥\bot), Moore-minimize (split blocks until every state of a block moves to the same block on every class), and number the blocks breadth-first from the start state taking bytes 0 to 255 in order. Two patterns with the same language get identical tables.

The JSON-schema subset of constrain.pyi maps to a regex over compact JSON: strings (well-formed UTF-8, the eight escapes, \uXXXX), integers, numbers, booleans, null, enum/const (each value as Python’s compact json.dumps, escaped), arrays with minItems/maxItems, objects whose properties appear in the order given, required ones always and optional ones possibly, and anyOf. Every other keyword is refused (422), never ignored. Property order is part of the language, so the schema is read with a JSON reader that keeps key order.

Token tt is allowed in state ss when δ∗(s,wt)≠⊥\delta^*(s, w_t) \ne \bot; EOS is allowed exactly when s∈As \in A; a special token or an empty token never. Walking every token’s bytes per state costs O(∑t∣wt∣)O(\sum_t |w_t|); a trie over the vocabulary shares the walk between tokens with a common prefix and prunes a whole subtree at the first ⊥\bot. Masks are cached per state and shared by every request using the same grammar. Sampling then runs L10.1’s sample on the logits with every disallowed token at −∞-\infty; the logprob is renormalized over the allowed tokens.

The model writes calls as <tool_call>{"name":"<name>","arguments":<object>}</tool_call>, the format small instruct models are trained on, separated by \n. tool_choice picks the grammar:

tool_choicegrammar
{"type":"function","function":{"name":"f"}}exactly one call of f, from the first token
"required"one call of any tool, or up to 4 with parallel_tool_calls
"auto" (default with tools)free text; once the output ends with <tool_call>, the rest of one call
"none"free text (or a JSON object with response_format: json_object, nested at most 3 levels: the language of all JSON objects is not regular)

The tools are described to the model in a system message (appended to the first system message if there is one); an assistant turn’s tool_calls are written back as <tool_call> blocks and a tool message as a <tool_response> user turn, so a conversation renders the way it was generated.

The parser keeps a small state machine: outside a call it emits text but holds back a suffix that could be the start of <tool_call>; after the opener it waits for the header {"name":"...","arguments": and emits the call’s id and name; inside the arguments it tracks brace depth and string escapes and passes every new piece on as an arguments fragment until the value closes; then it consumes }</tool_call>. Ids are call_<request id>_<index>. The stream’s chunks are {"tool_calls":[{"index":i,"id":...,"type":"function","function":{"name":...,"arguments":""}}]} for a call’s start and {"tool_calls":[{"index":i,"function":{"arguments":"<fragment>"}}]} after; finish_reason is tool_calls when any call was made.

The automaton (test hand_example_dfa). (cat|car|dog)s?: breadth-first from state 0 with bytes in order: c (0x63) gets id 1, d (0x64) id 2; from 1, a gives 3; from 2, o gives 4; from 3, r and t both reach the state after a whole word, and so does g from 4: minimization merges “car”, “cat”, and “dog” into state 5 (accepting); s from 5 gives 6 (accepting, no transitions). Seven states, accepting {5, 6}, eight live edges.

A mask (test constrained_sample_hand_example). Pattern ab?, vocabulary [a, b, c, EOS], logits [0, 1, 5, 2]. In state 0 only a walks live: greedy takes a although c has the largest logit, with logprob ln⁡1=0\ln 1 = 0. In the state after a (accepting) b and EOS are allowed: greedy takes EOS (2>12 > 1) with logprob 2−ln⁡(e1+e2)2 - \ln(e^1 + e^2).

A schema (grammar_by_tool_choice). The conformance tool’s parameters {"type":"object","properties":{"city":{"type":"string"}},"required":["city"]} become \{\"city\":"(?:...)*"\}: the escaped key, a colon, the JSON string pattern. Optional members nest: two optional integers x, y give \{(?:\"x\":INT(?:,\"y\":INT)?|(?:\"y\":INT|))\}, so {}, {"x":1}, {"y":2}, and {"x":1,"y":2} match and {"y":2,"x":1} does not.

A streamed call (test parser_hand_example). Pieces Sure.<tool_, call>{"name":"get_weather","argu, ments":{"city":"Pa, ris"}}</tool_call>:

PieceEvents
1content Sure. (<tool_ held back)
2nothing (the header is incomplete)
3call 0 starts (call_req7_0, get_weather); arguments {"city":"Pa
4arguments ris"}; }</tool_call> consumed

Folded: content Sure., one call with arguments {"city":"Paris"}, finish_reason tool_calls.

rust/crates/tl-engine/src/constrain.rs
pub const DEAD: i32 = -1;
pub struct Dfa { pub n_states: usize, pub accept: Vec<bool>, pub trans: Vec<i32> /* n * 256 */ }
impl Dfa { pub fn step(&self, state: i32, data: &[u8]) -> Result<i32, ConstrainError>; pub fn matches(&self, data: &[u8]) -> bool; }
pub fn parse(pattern: &str) -> Result<Node, ConstrainError>;
pub fn regex_to_dfa(pattern: &str) -> Result<Dfa, ConstrainError>;
pub enum Json { Null, Bool(bool), Int(String), Float(f64), Str(String), Arr(Vec<Json>), Obj(Vec<(String, Json)>) }
impl Json { pub fn parse(text: &str) -> Result<Json, ConstrainError>; pub fn get(&self, key: &str) -> Option<&Json>; pub fn dumps(&self) -> String; }
pub fn json_quote(s: &str) -> String; pub fn python_float(x: f64) -> String; pub fn escape(text: &str) -> String;
pub fn json_schema_to_regex(schema: &Json) -> Result<String, ConstrainError>;
pub fn json_object_regex(depth: usize) -> String;
pub struct TokenIndex { pub dfa: Dfa, pub vocab_size: usize, pub eos_id: Option<u32>, /* trie, cache */ }
impl TokenIndex { pub fn new(dfa: Dfa, vocab: &[Option<Vec<u8>>], eos_id: Option<u32>) -> Result<TokenIndex, ConstrainError>;
pub fn mask(&self, state: i32) -> Result<Vec<bool>, ConstrainError>; pub fn next_state(&self, state: i32, token_id: u32) -> Result<i32, ConstrainError>; }
pub struct Constraint { pub index: Arc<TokenIndex>, pub state: i32, pub done: bool } // new, mask, advance, is_complete
pub fn apply_mask(logits: &[f32], mask: &[bool]) -> Result<Vec<f32>, ConstrainError>;
pub fn constrained_sample(logits: &[f32], c: &mut Constraint, p: &SamplingParams, prompt: &[u32], output: &[u32], rng: &mut Pcg32) -> Result<(u32, f64), ConstrainError>;
// rust/crates/tl-serve/src/tools.rs
pub struct ToolError { pub status: u16, pub param: String, pub code: Option<String>, pub message: String } // to_api_error, to_json
pub struct Tool { pub name: String, pub description: Option<String>, pub parameters: Json }
pub enum ToolChoice { Auto, None, Required, Named(String) } pub enum ResponseFormat { Text, JsonObject }
pub struct ToolRequest { pub tools: Vec<Tool>, pub choice: ToolChoice, pub response_format: ResponseFormat, pub parallel: bool }
pub fn parse_tool_request(body: &str) -> Result<ToolRequest, ToolError>;
pub struct Grammar { pub pattern: String, pub trigger: Option<String> }
pub fn grammar(r: &ToolRequest) -> Option<Grammar>;
pub fn render_messages(messages: &Json, tools: &[Tool]) -> Result<Vec<(String, String)>, ToolError>;
pub fn render_prompt(t: &Template, messages: &Json, tools: &[Tool], bos: &str, eos: &str) -> Result<String, ToolError>;
pub fn vocab_bytes(tok: &dyn Tokenizer, specials: &[u32]) -> Vec<Option<Vec<u8>>>;
pub enum ToolEvent { Content(String), CallStart { index: usize, id: String, name: String }, Arguments { index: usize, fragment: String } }
impl ToolParser { pub fn new(request_id: &str) -> ToolParser; pub fn push(&mut self, piece: &str) -> Vec<ToolEvent>; pub fn finish(&mut self) -> Vec<ToolEvent>; }
pub fn collect(events: &[ToolEvent]) -> (Option<String>, Vec<ToolCall>);
pub fn finish_reason(calls: usize, engine: &str) -> String;
pub fn message_json(content: Option<&str>, calls: &[ToolCall]) -> String;
pub fn delta_json(e: &ToolEvent) -> String;
// rust/crates/tl-serve/src/server.rs, taken over from L10.7: its public API is unchanged
// (ServeConfig, spawn, run, ServerHandle); POST /v1/chat/completions accepts tools,
// tool_choice, response_format, and parallel_tool_calls.

course/fixtures/L10.9/constrain_golden.json (course/oracle/L10.9/constrain_golden.py) records, from the Python reference of L8.7, the regex and canonical DFA of 13 schemas and 8 patterns, the masks along three seeded paths of each through a 123-entry vocabulary (printable bytes, multi-byte and UTF-8 tokens, a special, an empty token, EOS), Python’s compact json.dumps of 19 JSON texts, and the schemas and patterns the subset refuses; L8.7’s own golden file (an independent derivative-based oracle) is replayed too.

TestKINDChecksWhy it matters downstream
hand_example_dfaunitsection 3’s seven states and eight edges; step and matchesyou and Python agree on the canonical form
dfa_matches_python_l8_7differentialregex text per schema; states, accepting flags, every transition per patternidentical tables, so identical masks
masks_match_python_l8_7differentialallowed ids and completeness at every step of every recorded pathwhat the sampler sees is Python’s
subset_errors_are_refusedboundary15 bad patterns and 11 bad schemas refuseda 422, never a silently different language
dumps_matches_pythonunitkey order, -0, float repr, escapes, raw non-ASCIIenum values the model can actually produce
token_index_rulesboundaryspecials, empty tokens, EOS rules, forbidden advancesthe edge cases of the mask
constrained_sample_hand_exampleunitsection 3’s mask and logprobs; an all-false mask refusedthe sampler sees the mask first
constrained_output_always_parsesproperty25 random walks per schema parse and validatethe guarantee itself
json_object_grammarboundaryobjects up to depth 3; arrays, scalars, broken objects refusedresponse_format: json_object
parse_tool_request_hand_exampleunitdefaults, a named choice, 400s and the 422 naming the keywordthe request surface of openai-subset.v1
grammar_by_tool_choiceunitnamed, required, parallel off, auto with its trigger, nonetools.choice
render_messages_with_toolsunittools in the system message; calls and results rendered back; through the chat templatemulti-turn agent loops (ag.03)
vocab_bytes_marks_specialsunitspecials map to Nonechat markers are never generated
parser_hand_exampleunitsection 3’s stream, fragment by fragmenttools.stream
parser_holds_back_only_real_markupboundarya<b, a trailing <tool, an opener without a headerno markup in content, no content lost
parser_parallel_callsunitindices, ids, the newline between callsparallel calls
parser_is_split_invariantpropertyevery 2-way split and 200 random splits fold to the same answertokens cut text anywhere
message_and_delta_json_shapesconformancethe message and both delta shapes, arguments as a stringwhat ag.01 and the openai client parse
forced_call_from_a_random_model_is_validproperty40 random walks under a named choice give one valid get_weather callMS-L10’s PR check on a tiny random model
serve_loop_routes_tool_callsconformancea live server: a forced call comes back as schema-valid tool_calls with finish_reason tool_calls, from the prompt render_prompt builds; streamed deltas reassemble to the same call; tool_choice none calls nothing; tool-field errors keep the OpenAI shapeag.01 and the tools.* conformance cases read exactly this
PitfallSymptomCaught by
Stopping minimization after one refinementequal languages, different tables: masks differ from Python’sdfa_matches_python_l8_7 (mutant s01)
Sorting properties by namethe model must write keys in an order the schema did not ask fordfa_matches_python_l8_7 (mutant s02)
No comma before a later optional member{"x":1"y":2} is “valid” and does not parseconstrained_output_always_parses (mutant s03)
Enum values inserted unescaped2.5 matches 2x5; quotes break the patterndfa_matches_python_l8_7 (mutant s04)
maxItems off by oneone item too many passes the grammar and fails the schemaconstrained_output_always_parses (mutant s05)
Walking a token’s later bytes from the wrong statemulti-byte tokens allowed that leave the languagemasks_match_python_l8_7 (mutant s06)
EOS allowed before a whole stringoutputs end mid-objecttoken_index_rules (mutant s07)
An empty language not refuseda request that can never finishsubset_errors_are_refused (mutant s08)
Floats printed unlike Python1.5e-05 written 0.000015: the enum value is unreachabledumps_matches_python (mutant s09)
Tokens still allowed after EOSgeneration continues past the end of the languagemasks_match_python_l8_7 (mutant s10)
The sampler seeing the unmasked logitsthe constraint is checked after the fact and failsconstrained_sample_hand_example (mutant s11)
Parameters outside the subset acceptedthe schema is silently weakenedparse_tool_request_hand_example (mutant s12)
A named tool_choice allowing every toolthe model calls the wrong functiongrammar_by_tool_choice (mutant s13)
Dropping argument fragmentsstreamed arguments arrive incompleteparser_hand_example (mutant s14)
Not holding back a partial <tool_call><tool_ shows up in the user-visible textparser_holds_back_only_real_markup (mutant s15)
content: "" next to tool_callsclients treat the turn as textmessage_and_delta_json_shapes (mutant s16)
DirectionModuleHow it uses this
BackL8.7the Python automaton and schema translation this port is held to
BackL10.1sample on the masked logits
BackL10.5Template::render_messages and ApiError
BackL1.5Tokenizer::token_bytes for the vocabulary’s bytes
BackL10.7the serve loop you take over: tool requests are recorded and traced through its calls
Forwardobs.02runs your engine, whose serve loop is now this module’s server.rs (the call site an upgrade inherits, DESIGN 3.4)
Forwardcraft.14takes your server.rs over to select API v1 or v2 per request (Pass 11)

ag.01 parses your tool_calls deltas; MS-agent runs SmolLM2-135M-Instruct through your engine with tool calls; the conformance cases tools.call, tools.stream, tools.choice turn from pending to checked once this module passes.

Your pieceProduction equivalentWhat it addsWhere to look
DFA masks cached per stateOutlinesthe same automaton approach, precomputed over the whole vocabularyOutlines
a regular subset of JSON schemaXGrammar, llguidancecontext-free grammars (recursive JSON) with a persistent stack and fast masksXGrammar, llguidance
<tool_call> text parsingvLLM tool parsersone parser per model family’s call formatvLLM tool calling
a lazy trigger in auto modellama.cpp grammar triggersgrammars activated by a token or pattern mid-generationllama.cpp grammars