In this article
Unstructured LLM outputs are a liability the moment an agent needs to take real actions. When a model free-forms its reasoning steps, a single hallucinated bracket or unexpected newline is enough to crash a parser, silently swallow a tool call, or—worse—send a malformed request to an external API. The ReAct framework (Reasoning + Acting) addresses the logic of interleaving thought and action, but it says nothing about format enforcement. Outlines fills that gap by intercepting the token-sampling process itself, making it structurally impossible for the model to emit output that violates a declared JSON schema. The result is an agent loop that is not just prompted to behave well, but mathematically constrained to do so.
This guide is aimed at engineers who are already comfortable with LLM APIs and want to graduate from "hope the model follows instructions" to "the model cannot do anything else." The technique applies equally to research agents that query knowledge bases, coding assistants that dispatch tool calls, or production pipelines where a downstream service deserves a strict contract. The example runs against gpt-4o-mini. The structured-generation overhead is CPU-bound index construction on your machine, which completes in seconds for the schemas used here. No GPU is required, and no fine-tuning is involved.
This guide is adapted from Outlines' react.py example, published under the Apache-2.0 licence. Code blocks below are reproduced exactly from that source.
Prerequisites
- Python 3.10+ with
outlines,openai, andrequestsinstalled (pip install outlines openai requests) - An OpenAI API key exported as
OPENAI_API_KEY;gpt-4o-miniis used throughout - Outlines ≥ 0.2 — the
Generator,Template, andJsonSchemaimports shown below reflect the current public API - Internet access for the live Wikipedia calls made during the agent loop
No local GPU, no Hugging Face account, and no vector database are required for this walkthrough.
Step 1: Define the prompt templates
ReAct agents are inherently multi-turn: each reasoning step appends to a growing context. Outlines' Template class wraps Jinja2 rendering, giving you a typed, reusable callable instead of raw f-strings. Two templates do all the prompt construction here.
import json
import requests # type: ignore
from openai import OpenAI
import outlines
from outlines import Generator, Template
from outlines.types import JsonSchema
build_reAct_prompt = Template.from_string(
"""What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into?
Tho 1: I need to search Colorado orogeny, find the area that the eastern sector of the Colorado ...
Act 2: Search 'Colorado orogeny'
Obs 2: The Colorado orogeny was an episode of mountain building (an orogeny) ...
Tho 3: It does not mention the eastern sector. So I need to look up eastern sector.
...
Tho 4: High Plains rise in elevation from around 1,800 to 7,000 ft, so the answer is 1,800 to 7,000 ft.
Act 5: Finish '1,800 to 7,000 ft'
{{ question }}
"""
)
add_mode = Template.from_string(
"""{{ prompt }}
{{ mode }} {{ i }}: {{ result }}
"""
)
build_reAct_prompt seeds the context with a worked example—a few-shot demonstration that teaches the model the Tho / Act / Obs labelling convention without requiring any fine-tuning. add_mode is then called on every loop iteration to append the current step's label and content. The separation matters: the few-shot prompt is written once, while add_mode handles the growing transcript.
Step 2: Wire up the model and define constrained generators
It is worth being precise about where the guarantee comes from, because it depends on the backend. This example builds its model with outlines.from_openai, and for OpenAI models Outlines does not touch token sampling: it converts each JsonSchema into OpenAI's structured-output request, a response_format of type json_schema, and OpenAI constrains the output server-side. What Outlines adds is the interface. The same Generator(model, mode_schema) call works unchanged against a local model loaded with outlines.from_transformers or outlines.from_vllm, and there Outlines enforces the schema itself by masking tokens that would break it during decoding. Either way, the narrow enum means result comes back as one of the declared values, which is what makes the routing below safe.
def search_wikipedia(query: str):
url = f"https://en.wikipedia.org/w/api.php?format=json&action=query&prop=extracts&exintro&explaintext&redirects=1&titles={query}&origin=*"
response = requests.get(url)
page = response.json()["query"]["pages"]
return ".".join(list(page.values())[0]["extract"].split(".")[:2])
prompt = build_reAct_prompt(question="Where is Apple Computers headquarted? ")
model = outlines.from_openai(OpenAI(), "gpt-4o-mini")
# Define JSON schemas for mode and action
mode_schema = JsonSchema({
"type": "object",
"properties": {
"result": {
"type": "string",
"enum": ["Tho", "Act"]
}
},
"required": ["result"]
})
action_schema = JsonSchema({
"type": "object",
"properties": {
"result": {
"type": "string",
"enum": ["Search", "Finish"]
}
},
"required": ["result"]
})
mode_generator = Generator(model, mode_schema)
action_generator = Generator(model, action_schema)
text_generator = Generator(model)
Three generators are constructed from the same underlying model, each with a different output contract. mode_generator can only ever return {"result": "Tho"} or {"result": "Act"}. action_generator is similarly restricted to Search or Finish. text_generator is unconstrained and is used for the freeform content within each step—the actual thought text or the Wikipedia search subject. The design deliberately keeps structure enforcement narrow: only the routing decisions are locked down; the content remains flexible.
Step 3: Run the structured agent loop
The loop is the heart of the implementation. Each iteration asks the constrained mode_generator to classify what kind of step comes next, branches on the result, and either generates a thought or dispatches a tool call.
for i in range(1, 10):
mode_output = mode_generator(prompt, max_tokens=128)
mode = json.loads(mode_output)["result"] # Extract the result from the JSON output
prompt = add_mode(i=i, mode=mode, result="", prompt=prompt)
if mode == "Tho":
thought = text_generator(prompt, stop="\n", max_tokens=128)
prompt += f"{thought}"
elif mode == "Act":
action_output = action_generator(prompt, max_tokens=128)
action = json.loads(action_output)["result"] # Extract the result from the JSON output
prompt += f"{action} '"
subject = text_generator(prompt, stop=["'"], max_tokens=128)
# Apple Computers headquartered
subject = " ".join(subject.split()[:2])
prompt += f"{subject}'"
if action == "Search":
result = search_wikipedia(subject)
prompt = add_mode(i=i, mode="Obs", result=result, prompt=prompt)
else:
break
print(prompt)
Notice the pattern: a constrained generator decides what kind of token sequence is valid next, and an unconstrained generator fills in the value. The stop="\n" argument on the thought generator prevents the model from writing multiple lines when only one is expected. When the action is Search, the Wikipedia result is injected as an Obs line, extending the transcript so the next iteration has fresh context. When the action is Finish, the loop breaks cleanly—no try/except needed because the model cannot emit an unrecognised action token.
Step 4: Inspect the printed transcript
After the loop, print(prompt) reveals the full agent transcript: every thought, action, observation, and the final answer, formatted exactly as the few-shot preamble taught the model to write. A successful run looks roughly like this in the terminal:
...
Tho 3: I need to find where Apple Computers is headquartered.
Act 3: Search 'Apple Computers'
Obs 3: Apple Inc. is an American multinational technology company headquartered in Cupertino, California.
Tho 4: Apple Computers is headquartered in Cupertino, California.
Act 4: Finish 'Cupertino, California'
The labels themselves are written by the loop, not the model: add_mode() appends Tho i:, Act i: or Obs i: using the value the constrained generator returned, so the transcript format cannot drift. Debugging is straightforward: if something goes wrong, the transcript is the complete reproducible record of what each generator decided and why.
Generator and schema options at a glance
| Generator instance | Schema / constraint | Allowed values | Role in the loop |
|---|---|---|---|
mode_generator |
mode_schema (enum) |
Tho, Act |
Routing: determines the step type |
action_generator |
action_schema (enum) |
Search, Finish |
Tool dispatch: picks which tool or terminates |
text_generator |
None | Arbitrary string | Content: thought text or search subject |
What to watch out for
Search subjects are cut to two words. After the model writes the search subject, the loop keeps only its first two words (subject = " ".join(subject.split()[:2])) before calling Wikipedia. That is why the example question works, since "Apple Computers" survives intact, but a three-word subject such as "Colorado orogeny eastern" is silently shortened. Remove or widen that line before trying longer queries.
max_tokens on constrained generators is still a hard ceiling. The structured output guarantee does not waive the token budget. If your JSON schema requires more tokens to close properly than max_tokens allows, the output will be truncated mid-JSON, and json.loads will throw. Keep schema outputs small, or raise the budget deliberately.
Wikipedia's extracts API returns raw wikitext artifacts. The .split(".")[:2] slicing in search_wikipedia is a rough heuristic—it keeps only the first two sentences. For topics with complex lead sections this can cut mid-fact. In production you would want a cleaner extraction step, especially if the observation feeds into a downstream system.
The subject = " ".join(subject.split()[:2]) truncation is intentional but lossy. Multi-word proper nouns like "Cupertino, California" or "New York City" get reduced to two tokens before the Wikipedia query. This works for the example, but it is a silent correctness hazard for any entity whose canonical Wikipedia title is longer. Consider a named-entity extraction pass or a fuzzy-search endpoint if your domain has complex entity names.
The loop has no backtracking. If the model emits a thought that contradicts the observation, there is no mechanism to revise it. The transcript is append-only. For agents that need to recover from bad intermediate steps, you would need to checkpoint the prompt at each observation and allow rollback—a meaningful architectural addition that this example does not include.
Security: agent tool calls and external data. Injecting raw Wikipedia text into the prompt without sanitisation creates a prompt-injection surface. An adversarial Wikipedia article could in principle steer the agent toward a harmful Finish value. The constraints on action_schema limit the blast radius, but they do not eliminate it. For a fuller discussion of agent security exposure, see our coverage of the OpenAI agent that inadvertently posted user images publicly and the health portal disclosure involving an OpenAI agent—both illustrate what can go wrong when agent I/O boundaries are not treated as a trust boundary.
Where to go next
The architecture here is intentionally minimal: one model, two constrained generators, one external tool. Natural extensions include adding more entries to action_schema—Calculate, LookupDatabase, CallAPI—each paired with its own tool function and result injection. You can also replace the Wikipedia call with any deterministic service that returns a string, making the pattern generic.
For tighter latency budgets, consider replacing gpt-4o-mini with a locally-hosted model via outlines.from_transformers or outlines.from_vllm. The constrained generation machinery is model-agnostic; only the model = outlines.from_openai(...) line changes. If you need to reduce the number of tokens the model spends on reasoning before reaching Finish, techniques like those explored in ThinkingCap's approach to fewer thinking tokens offer complementary strategies worth reviewing alongside this one.
The full source is at github.com/dottxt-ai/outlines under the Apache-2.0 licence. The project's documentation covers regex-constrained generation, grammar-constrained generation, and integration with vLLM for high-throughput serving—all of which compose cleanly with the ReAct pattern shown here.
Frequently asked questions
What is Outlines and how does it enforce JSON output?
Outlines (github.com/dottxt-ai/outlines, Apache-2.0) gives every backend one interface for structured generation, and enforces the schema in whatever way that backend supports. With a local backend such as from_transformers or from_vllm it masks invalid tokens during decoding itself. With from_openai, which this example uses, it converts the JsonSchema into OpenAI's structured-output request (response_format with a json_schema), and OpenAI enforces it server-side.
Does Outlines constrained generation work with OpenAI models?
Yes. `outlines.from_openai(OpenAI(), "gpt-4o-mini")` wraps any OpenAI-compatible endpoint, and the same `Generator` API works with local models via `outlines.from_transformers` or `outlines.from_vllm`. Only the model-construction line changes; the schema and generator code stays identical.
Why use three separate generators instead of one?
Each generator carries a different output contract. `mode_generator` is locked to `{"result": "Tho" | "Act"}`, `action_generator` to `{"result": "Search" | "Finish"}`, and `text_generator` is unconstrained for freeform content. Keeping constraints narrow means routing decisions are guaranteed correct while thought text and search subjects remain flexible.
How much does it cost to run this example?
The example targets `gpt-4o-mini`, which is among the cheapest OpenAI tiers. A full debugging session—multiple loop runs of up to 10 iterations each—typically costs a few cents. The constrained-generation overhead is CPU-bound index construction on your machine and completes in seconds for the small enum schemas used here.
What happens if max_tokens is too low for the constrained generator?
The structured output guarantee does not override the token budget. If the schema requires more tokens to close than `max_tokens` allows, the output is truncated mid-JSON and `json.loads` will raise a decode error. Keep schema outputs small or raise the budget deliberately when using larger schemas.
Can I add more tools to this agent pattern?
Yes. Add new string values to the `action_schema` enum—for example `"Calculate"` or `"CallAPI"`—then add a matching branch inside the `elif mode == "Act"` block that calls the corresponding function and injects the result as an `Obs` line. The constrained generator will automatically restrict the model to only the actions you have declared.
Related Guides

Generate and Run a Python Expression with Outlines
Use Outlines and GPT-4o mini to turn a word problem into a Python expression, execute it, and understand the risks of eval.

Self-Consistency Voting with Outlines and gpt-4o-mini
Generate ten reasoning chains in one API call, extract integer answers with regex, and vote for the majority—reliably solving multi-step arithmetic.

Build a Text-to-SQL Agent with smolagents in One File
Wire a Hugging Face CodeAgent to a SQLite database so it answers plain-English questions with verified SQL queries.