In this article
Most LLM demonstrations stop after generating text. This example runs the generated text: a hosted model converts a word problem into one Python expression, which the script executes locally with eval.
The implementation uses Outlines for prompt templating and model access, OpenAI’s gpt-4o-mini, and three few-shot examples. It is compact, but its unconstrained output and use of eval make it suitable as a learning example rather than a safe service pattern.
Adapted from Outlines’ math_generate_code.py example. Outlines is distributed under the Apache-2.0 licence.
Prerequisites
Install the two imported packages in your Python environment:
pip install openai outlines
Set an OpenAI API key in the environment used to run the script:
export OPENAI_API_KEY="your-api-key"
On Windows PowerShell, use:
$env:OPENAI_API_KEY="your-api-key"
The example sends its prompt to the hosted gpt-4o-mini model through the OpenAI client.
The complete source
The code below matches the tested source, including the original question wording and introductory source comment:
"""Example from https://dust.tt/spolu/a/d12ac33169"""
import openai
import outlines
from outlines import Template
examples = [
{"question": "What is 37593 * 67?", "code": "37593 * 67"},
{
"question": "Janet's ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market?",
"code": "(16-3-4)*2",
},
{
"question": "A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take?",
"code": " 2 + 2/2",
},
]
question = "Carla is downloading a 200 GB file. She can download 2 GB/minute, but 40% of the way through the download, the download fails. Then Carla has to restart the download from the beginning. How load did it take her to download the file in minutes?"
answer_with_code_prompt = Template.from_string(
"""
{% for example in examples %}
QUESTION: {{example.question}}
CODE: {{example.code}}
{% endfor %}
QUESTION: {{question}}
CODE:"""
)
def execute_code(code):
result = eval(code)
return result
prompt = answer_with_code_prompt(question=question, examples=examples)
model = outlines.from_openai(openai.OpenAI(), "gpt-4o-mini")
answer = model(prompt)
result = execute_code(answer)
print(f"It takes Carla {result:.0f} minutes to download the file.")
1. Define the few-shot examples
The examples list pairs arithmetic questions with one-line Python expressions:
examples = [
{"question": "What is 37593 * 67?", "code": "37593 * 67"},
{
"question": "Janet's ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market?",
"code": "(16-3-4)*2",
},
{
"question": "A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take?",
"code": " 2 + 2/2",
},
]
These demonstrations establish the desired mapping:
| Input | Expected output form |
|---|---|
| Arithmetic question | A Python expression such as 37593 * 67 |
| Word problem | An expression combining the quantities and operators |
| Model response | Expression only, without an explanation or Markdown fence |
The examples do not define functions, import modules, or ask the model for a full script. They demonstrate only expression generation, although the script does not technically enforce that restriction.
2. Build the prompt template
The target question describes a failed download followed by a complete restart:
question = "Carla is downloading a 200 GB file. She can download 2 GB/minute, but 40% of the way through the download, the download fails. Then Carla has to restart the download from the beginning. How load did it take her to download the file in minutes?"
The source contains the phrase “How load” rather than “How long.” It is preserved here because the code must match the tested source.
Outlines’ Template loops over the examples and renders each pair under QUESTION: and CODE: labels:
answer_with_code_prompt = Template.from_string(
"""
{% for example in examples %}
QUESTION: {{example.question}}
CODE: {{example.code}}
{% endfor %}
QUESTION: {{question}}
CODE:"""
)
The final CODE: label is deliberately incomplete. It asks the model to continue the demonstrated pattern with an expression for the new question.
3. Render the prompt and call GPT-4o mini
The template becomes a concrete prompt when called with question and examples:
prompt = answer_with_code_prompt(question=question, examples=examples)
The script then wraps a synchronous OpenAI client with Outlines and selects gpt-4o-mini:
model = outlines.from_openai(openai.OpenAI(), "gpt-4o-mini")
answer = model(prompt)
answer is expected to be a string containing a Python arithmetic expression. This call is unconstrained: the model can still return prose, a CODE: prefix, Markdown fences, or invalid Python.
4. Execute and format the answer
The helper passes the generated string directly to Python’s eval:
def execute_code(code):
result = eval(code)
return result
The returned value is formatted with no digits after the decimal point:
result = execute_code(answer)
print(f"It takes Carla {result:.0f} minutes to download the file.")
The intended calculation is:
- Forty percent of 200 GB is 80 GB.
- At 2 GB per minute, the failed partial download takes 40 minutes.
- Restarting and downloading all 200 GB takes 100 minutes.
- The total is 140 minutes.
A valid generated expression could therefore be:
(200 * 0.4 / 2) + (200 / 2)
If the model returns an equivalent numeric expression, the final line prints:
It takes Carla 140 minutes to download the file.
Failure modes
The script assumes that the model returns both valid Python and a numeric result. Different outputs fail at different stages:
| Model output | Likely result | Reason |
|---|---|---|
(200 * 0.4 / 2) + (200 / 2) | Prints 140 minutes | Valid expression with a numeric result |
140 minutes | SyntaxError | The unit is not valid Python syntax |
| A fenced Markdown code block | SyntaxError | Backticks are passed to eval |
CODE: 140 | SyntaxError | The label is not part of a Python expression |
"140" | Formatting failure | The expression returns a string rather than a number |
2 / 0 | ZeroDivisionError | The expression is syntactically valid but fails during execution |
Structural validity would not establish that the arithmetic is correct. A model can generate a valid expression that represents the word problem incorrectly.
Do not expose this eval pattern to untrusted input
eval executes a string as a Python expression in the current process. The few-shot prompt asks for arithmetic, but that request does not prevent the model from returning expressions that access names, call functions, import modules, or trigger side effects.
For an application rather than a controlled demonstration:
- Do not pass unrestricted model output to
eval. - Parse and validate an intentionally narrow expression language.
- Permit only the numeric literals and arithmetic operators the application needs.
- Reject names, attribute access, indexing, function calls, imports, and other Python constructs.
- Handle malformed expressions, division by zero, API failures, and non-numeric results explicitly.
- Validate the computed answer independently when correctness matters.
Constrained generation can reduce format drift, but it does not make arbitrary Python execution safe and does not guarantee correct reasoning.
Choosing an output strategy
| Approach | What it controls | Main limitation | Appropriate use |
|---|---|---|---|
| Few-shot text generation | Demonstrates the desired response pattern | The model can ignore or drift from that pattern | Small experiments and inspected outputs |
| Constrained expression generation | Restricts the permitted output syntax | A valid expression can still encode incorrect arithmetic | Narrow expression languages with separate validation |
| Structured fields | Separates values into a defined data shape | Schema validity does not prove semantic correctness | Tool arguments, classifications, and typed application data |
The tested code follows the first approach. Its useful lesson is the complete flow from examples, to a rendered prompt, to a model-generated expression, to a computed result. Its equally important lesson is that generated code needs a validation boundary before execution.
For another example of deterministic controls around AI-generated code, read Alibaba’s deterministic AI code review overview.
Frequently asked questions
How do I use Outlines with an OpenAI model?
Create an OpenAI client, then pass it and the model name to `outlines.from_openai`. In this example, `outlines.from_openai(openai.OpenAI(), "gpt-4o-mini")` returns a model handle that accepts the rendered prompt.
What does the Outlines Template class do?
`Template.from_string` creates a reusable prompt template with Jinja-style variables and loops. This example inserts three question-and-code demonstrations before appending the new word problem and an incomplete `CODE:` line.
Why is using eval on LLM output dangerous?
`eval` executes the supplied string as Python rather than treating it as inert text. Model output can therefore access names, call functions, import modules, or perform other unintended operations, so this pattern must not be exposed to untrusted input.
What answer should the Carla download example produce?
Carla spends 40 minutes downloading 40% of the 200 GB file at 2 GB per minute. After the failure, the complete restart takes another 100 minutes, for a total of 140 minutes.
Does this Outlines example guarantee a valid Python expression?
No. The tested source uses an unconstrained model call, so the few-shot prompt is the only instruction enforcing the output format. Prose, Markdown fences, units, or invalid Python can cause `eval` or the numeric format operation to fail.
Related Guides

Self-Consistency Voting with Outlines and gpt-4o-mini
Generate ten reasoning chains in one API call, extract integer answers with regex, and vote for the majority—reliably solving multi-step arithmetic.

Build a Deep Research API Agent Pipeline with OpenAI
Run a four-agent Deep Research pipeline—clarification, prompt enrichment, MCP file search, and citation extraction—using OpenAI's agents SDK.
Mastering Advanced RAG Evaluation: From Basic Metrics to LLM-as-a-Judge
Learn how to systematically evaluate Retrieval-Augmented Generation (RAG) pipelines using the RAG Triad, Ragas, and LLM-as-a-Judge techniques.