Generate and Run a Python Expression with Outlines

October 2, 2026 • guides
OpenAIPythonLLMs

Most LLM demonstrations stop after generating text. This example runs the generated text: a hosted model converts a word problem into one Python expression, which the script executes locally with eval.

The implementation uses Outlines for prompt templating and model access, OpenAI’s gpt-4o-mini, and three few-shot examples. It is compact, but its unconstrained output and use of eval make it suitable as a learning example rather than a safe service pattern.

Adapted from Outlines’ math_generate_code.py example. Outlines is distributed under the Apache-2.0 licence.

Prerequisites

Install the two imported packages in your Python environment:

pip install openai outlines

Set an OpenAI API key in the environment used to run the script:

export OPENAI_API_KEY="your-api-key"

On Windows PowerShell, use:

$env:OPENAI_API_KEY="your-api-key"

The example sends its prompt to the hosted gpt-4o-mini model through the OpenAI client.

The complete source

The code below matches the tested source, including the original question wording and introductory source comment:

"""Example from https://dust.tt/spolu/a/d12ac33169"""

import openai

import outlines
from outlines import Template


examples = [
    {"question": "What is 37593 * 67?", "code": "37593 * 67"},
    {
        "question": "Janet's ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market?",
        "code": "(16-3-4)*2",
    },
    {
        "question": "A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take?",
        "code": " 2 + 2/2",
    },
]

question = "Carla is downloading a 200 GB file. She can download 2 GB/minute, but 40% of the way through the download, the download fails. Then Carla has to restart the download from the beginning. How load did it take her to download the file in minutes?"

answer_with_code_prompt = Template.from_string(
    """
    {% for example in examples %}
    QUESTION: {{example.question}}
    CODE: {{example.code}}

    {% endfor %}
    QUESTION: {{question}}
    CODE:"""
)


def execute_code(code):
    result = eval(code)
    return result


prompt = answer_with_code_prompt(question=question, examples=examples)
model = outlines.from_openai(openai.OpenAI(), "gpt-4o-mini")
answer = model(prompt)
result = execute_code(answer)
print(f"It takes Carla {result:.0f} minutes to download the file.")

1. Define the few-shot examples

The examples list pairs arithmetic questions with one-line Python expressions:

examples = [
    {"question": "What is 37593 * 67?", "code": "37593 * 67"},
    {
        "question": "Janet's ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market?",
        "code": "(16-3-4)*2",
    },
    {
        "question": "A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take?",
        "code": " 2 + 2/2",
    },
]

These demonstrations establish the desired mapping:

InputExpected output form
Arithmetic questionA Python expression such as 37593 * 67
Word problemAn expression combining the quantities and operators
Model responseExpression only, without an explanation or Markdown fence

The examples do not define functions, import modules, or ask the model for a full script. They demonstrate only expression generation, although the script does not technically enforce that restriction.

2. Build the prompt template

The target question describes a failed download followed by a complete restart:

question = "Carla is downloading a 200 GB file. She can download 2 GB/minute, but 40% of the way through the download, the download fails. Then Carla has to restart the download from the beginning. How load did it take her to download the file in minutes?"

The source contains the phrase “How load” rather than “How long.” It is preserved here because the code must match the tested source.

Outlines’ Template loops over the examples and renders each pair under QUESTION: and CODE: labels:

answer_with_code_prompt = Template.from_string(
    """
    {% for example in examples %}
    QUESTION: {{example.question}}
    CODE: {{example.code}}

    {% endfor %}
    QUESTION: {{question}}
    CODE:"""
)

The final CODE: label is deliberately incomplete. It asks the model to continue the demonstrated pattern with an expression for the new question.

3. Render the prompt and call GPT-4o mini

The template becomes a concrete prompt when called with question and examples:

prompt = answer_with_code_prompt(question=question, examples=examples)

The script then wraps a synchronous OpenAI client with Outlines and selects gpt-4o-mini:

model = outlines.from_openai(openai.OpenAI(), "gpt-4o-mini")
answer = model(prompt)

answer is expected to be a string containing a Python arithmetic expression. This call is unconstrained: the model can still return prose, a CODE: prefix, Markdown fences, or invalid Python.

4. Execute and format the answer

The helper passes the generated string directly to Python’s eval:

def execute_code(code):
    result = eval(code)
    return result

The returned value is formatted with no digits after the decimal point:

result = execute_code(answer)
print(f"It takes Carla {result:.0f} minutes to download the file.")

The intended calculation is:

  • Forty percent of 200 GB is 80 GB.
  • At 2 GB per minute, the failed partial download takes 40 minutes.
  • Restarting and downloading all 200 GB takes 100 minutes.
  • The total is 140 minutes.

A valid generated expression could therefore be:

(200 * 0.4 / 2) + (200 / 2)

If the model returns an equivalent numeric expression, the final line prints:

It takes Carla 140 minutes to download the file.

Failure modes

The script assumes that the model returns both valid Python and a numeric result. Different outputs fail at different stages:

Model outputLikely resultReason
(200 * 0.4 / 2) + (200 / 2)Prints 140 minutesValid expression with a numeric result
140 minutesSyntaxErrorThe unit is not valid Python syntax
A fenced Markdown code blockSyntaxErrorBackticks are passed to eval
CODE: 140SyntaxErrorThe label is not part of a Python expression
"140"Formatting failureThe expression returns a string rather than a number
2 / 0ZeroDivisionErrorThe expression is syntactically valid but fails during execution

Structural validity would not establish that the arithmetic is correct. A model can generate a valid expression that represents the word problem incorrectly.

Do not expose this eval pattern to untrusted input

eval executes a string as a Python expression in the current process. The few-shot prompt asks for arithmetic, but that request does not prevent the model from returning expressions that access names, call functions, import modules, or trigger side effects.

For an application rather than a controlled demonstration:

  1. Do not pass unrestricted model output to eval.
  2. Parse and validate an intentionally narrow expression language.
  3. Permit only the numeric literals and arithmetic operators the application needs.
  4. Reject names, attribute access, indexing, function calls, imports, and other Python constructs.
  5. Handle malformed expressions, division by zero, API failures, and non-numeric results explicitly.
  6. Validate the computed answer independently when correctness matters.

Constrained generation can reduce format drift, but it does not make arbitrary Python execution safe and does not guarantee correct reasoning.

Choosing an output strategy

ApproachWhat it controlsMain limitationAppropriate use
Few-shot text generationDemonstrates the desired response patternThe model can ignore or drift from that patternSmall experiments and inspected outputs
Constrained expression generationRestricts the permitted output syntaxA valid expression can still encode incorrect arithmeticNarrow expression languages with separate validation
Structured fieldsSeparates values into a defined data shapeSchema validity does not prove semantic correctnessTool arguments, classifications, and typed application data

The tested code follows the first approach. Its useful lesson is the complete flow from examples, to a rendered prompt, to a model-generated expression, to a computed result. Its equally important lesson is that generated code needs a validation boundary before execution.

For another example of deterministic controls around AI-generated code, read Alibaba’s deterministic AI code review overview.

Frequently asked questions

How do I use Outlines with an OpenAI model?

Create an OpenAI client, then pass it and the model name to `outlines.from_openai`. In this example, `outlines.from_openai(openai.OpenAI(), "gpt-4o-mini")` returns a model handle that accepts the rendered prompt.

What does the Outlines Template class do?

`Template.from_string` creates a reusable prompt template with Jinja-style variables and loops. This example inserts three question-and-code demonstrations before appending the new word problem and an incomplete `CODE:` line.

Why is using eval on LLM output dangerous?

`eval` executes the supplied string as Python rather than treating it as inert text. Model output can therefore access names, call functions, import modules, or perform other unintended operations, so this pattern must not be exposed to untrusted input.

What answer should the Carla download example produce?

Carla spends 40 minutes downloading 40% of the 200 GB file at 2 GB per minute. After the failure, the complete restart takes another 100 minutes, for a total of 140 minutes.

Does this Outlines example guarantee a valid Python expression?

No. The tested source uses an unconstrained model call, so the few-shot prompt is the only instruction enforcing the output format. Prose, Markdown fences, units, or invalid Python can cause `eval` or the numeric format operation to fail.

Free interactive tools for the decisions this piece raises.

Related Guides