All articlesThe Chronicle / Geodd

What Is OpenAI Structured Output? Schemas, Pydantic & Zod

You ask a model for JSON, and most of the time you get it. Then one response arrives with a missing field, a trailing comment, or a string where you expected an integer, and your parser throws. In production, that one failure can break an agent run. OpenAI structured output exists to remove that class of bug.

Here is the short answer. Structured Outputs is an API feature that forces the model to return JSON that matches a JSON Schema you supply. It goes beyond JSON mode, which only guarantees valid JSON. With a schema, you get the required keys, the right types, and enum values you defined. You can set it up with raw JSON Schema, Pydantic in Python, or Zod in TypeScript, in both the Responses API and Chat Completions.

Below, you will see how the feature works, how to define schemas in each format, and where it breaks down. We also cover how to use the same schema code with other providers. At Geodd, our OpenAI-compatible API lets you switch providers by changing one line, so this knowledge carries over.

Why structured outputs matter for production apps

What goes wrong without a schema

Prompting a model to "reply in JSON" is a request, not a contract. The model can drift from your format on long outputs, odd inputs, or a prompt change you made last Tuesday. Before schema enforcement, teams papered over this with regex cleanup and retry loops. OpenAI's own numbers show the gap: on its complex JSON schema eval, gpt-4o-2024-08-06 with Structured Outputs scored 100%, while gpt-4-0613 scored under 40%. That is the difference between a feature and a pager alert.

The failures are boring and repetitive, which is exactly why they hurt at scale.

FailureWhat breaks downstreamWhat a schema prevents
Missing required keyKeyError or null referenceAll required fields are present
Wrong type ("12" instead of 12)Database insert fails, math goes wrongTypes are enforced
Invented enum valueRouter finds no handlerOnly your listed values appear
Prose wrapped around the JSONjson.loads raisesOutput is the JSON and nothing else

What the guarantee covers

With OpenAI structured output turned on, the API constrains generation so the model cannot emit tokens that would violate your schema. You get valid JSON that matches the schema you supplied, with no format retry. Your code can parse the response and trust the shape of it. That is why a typed object in Pydantic or Zod comes back populated, not half-filled.

There is a limit, though. The guarantee is about structure, not truth. A model can return a perfectly typed invoice with a wrong total, or pick a valid enum value that is the wrong category. You still need business validation and evals on top.

A schema guarantees the shape of the answer, not the correctness of it.

Where it pays off in real systems

Agents feel this most. A multi-step agent chains model calls, and each output feeds the next step. One malformed response early in the chain can poison everything after it, and retrying a long run costs real tokens and minutes. Schema-bound output removes the most common cause of those restarts.

These are the places where teams get the biggest return:

  • Data extraction: pull invoice lines, contact details, or clinical codes from messy text into a fixed record.
  • Classification and routing: return one label from a closed list so your router never sees an unknown value.
  • Tool and agent steps: pass arguments between steps as typed objects instead of parsing prose.
  • UI generation: produce component props or form definitions your frontend can render directly.

Cost and latency matter too. Every retry doubles the spend for that request and adds a full round trip, so removing format failures cuts both tail latency and token waste. It also simplifies your code. You delete the cleanup layer, and your types become the single source of truth for what the model must return.

Finally, the idea travels. A schema is plain JSON Schema, so the same definitions work across LLM structured output setups that follow the OpenAI format. Support still varies by model and provider, so test your schema against the exact model you plan to ship.

How to use structured outputs in the Responses API

Where the schema goes

The OpenAI Responses API structured output setup differs from older code in one place. You pass the schema in text.format, not in response_format. Set type to json_schema, give the schema a name, add "strict": true, and include your JSON Schema. Without strict mode, the model treats the schema as a hint instead of a rule.

In the Responses API, the schema lives in text.format, and strict: true is what turns it into a guarantee.

from openai import OpenAI
client = OpenAI()

response = client.responses.create(
    model="gpt-4o-2024-08-06",
    input="Classify: 'My card was charged twice.'",
    text={"format": {
        "type": "json_schema",
        "name": "ticket",
        "strict": True,
        "schema": {
            "type": "object",
            "properties": {
                "category": {"type": "string", "enum": ["billing", "bug", "other"]},
                "urgent": {"type": "boolean"}
            },
            "required": ["category", "urgent"],
            "additionalProperties": False
        }
    }},
)

Strict mode demands that every property appears in required and that additionalProperties is false. Skip either rule and the API rejects the request with an error.

Reading the result

Next, read response.output_text. It is a JSON string, so one json.loads call gives you your object. Before you trust it, check three things:

  • response.status is completed, not incomplete. A low max_output_tokens can cut the JSON off mid-object.
  • The output contains no item of type refusal. The model returns one when it declines a request, and it will not match your schema.
  • Your own business checks pass, because the schema only covers shape.

Skipping manual parsing with the SDK

Finally, the official SDKs can do the parsing for you. In Python, client.responses.parse() accepts a Pydantic class through text_format and returns response.output_parsed as a typed object. The TypeScript SDK offers the same flow with Zod. You write the model once, and the SDK generates the JSON Schema and validates the reply. The next sections cover Chat Completions and the schema formats in detail.

How to use structured outputs in Chat Completions

Plenty of codebases still call chat.completions.create, and you do not need to rewrite them. OpenAI chat completion structured output follows the same strict schema rules as the Responses API. Only the place where you put the schema changes.

Moving the schema into response_format

In this API, the schema goes in response_format. Set type to json_schema, then nest a json_schema object holding the name, strict, and schema keys. That extra nesting level is the most common copy-paste mistake when you port code from the Responses API. The schema body stays identical, including required and additionalProperties: false.

response = client.chat.completions.create(
    model="gpt-4o-2024-08-06",
    messages=[{"role": "user", "content": "Classify: 'My card was charged twice.'"}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "ticket",
            "strict": True,
            "schema": ticket_schema,  # same schema as before
        },
    },
)

Reading content, refusals and truncation

The JSON string comes back in response.choices[0].message.content. Run json.loads on it, but check these fields first:

  • message.refusal holds text when the model declines. In that case content is empty or null.
  • finish_reason should be stop. A value of length means the output hit max_tokens and the JSON is cut off.
  • Your own business rules still apply, since the schema only covers shape.

Always check refusal and finish_reason before you parse, because a schema cannot protect you from an answer that never finished or never came.

Letting the SDK parse for you

Pydantic users can skip the manual steps. Call client.chat.completions.parse() and pass your class as response_format. You get message.parsed as a typed object. In older SDK versions, this method lived under client.beta.chat.completions. The TypeScript SDK does the same with Zod through its zodResponseFormat helper.

Because a schema is just a request field, this style of GPT API structured output also ports well. Many OpenAI-compatible providers accept the same response_format payload, so your code often needs only a new base URL and model name. Still, run your schema against the exact model you plan to ship, since strict enforcement is not universal.

Defining schemas with Pydantic, Zod and raw JSON Schema

You can describe one contract in three ways, and all three send the same strict JSON Schema to the API. Pick the one that fits your stack. Typed classes keep validation and schema in one place, while raw JSON Schema works in any language.

Pydantic in Python

Python teams should start with Pydantic. You define a class, pass it as text_format, and the SDK builds the schema and parses the reply for you.

from pydantic import BaseModel
from typing import Literal

class Ticket(BaseModel):
    category: Literal["billing", "bug", "other"]
    urgent: bool

response = client.responses.parse(
    model="gpt-4o-2024-08-06",
    input="Classify: 'My card was charged twice.'",
    text_format=Ticket,
)
ticket = response.output_parsed

Write the schema once as a typed class, and let the SDK generate the JSON Schema.

Optional fields need care. Strict mode requires every property in required, so use a nullable type such as str | None instead of a default value. The model then returns null when it has nothing to say.

Zod in TypeScript

TypeScript developers get the same flow with Zod. For OpenAI structured output with Zod, import zodTextFormat for the Responses API, or zodResponseFormat for Chat Completions.

import { z } from "zod";
import { zodTextFormat } from "openai/helpers/zod";

const Ticket = z.object({
  category: z.enum(["billing", "bug", "other"]),
  urgent: z.boolean(),
});

const response = await client.responses.parse({
  model: "gpt-4o-2024-08-06",
  input: "Classify: 'My card was charged twice.'",
  text: { format: zodTextFormat(Ticket, "ticket") },
});

The same rule applies here. Swap .optional() for .nullable(), because an optional key would break the all-fields-required rule.

Raw JSON Schema

Raw schemas make sense when you work in Go, Java, or Rust, or when you store schemas in config files. You write the object by hand, as in the earlier examples, and you own the strict-mode rules yourself. For OpenAI structured output JSON Schema work, that means required on every key and additionalProperties: false on every object, including nested ones.

OptionBest forOptional fields
PydanticPython servicesType | None
ZodTypeScript apps.nullable()
Raw JSON SchemaOther languages, shared config"type": ["string", "null"]

Whichever you choose, print the generated schema once and read it. A quick look catches a missing enum or an unexpected nested object before it costs you a failed request.

Structured outputs vs JSON mode vs function calling

Three features get mixed up here, and the mix-up costs teams time. Each one solves a different problem, so picking the wrong one leaves you with either weak guarantees or awkward code.

JSON mode guarantees syntax only

JSON mode, set with {"type": "json_object"}, makes sure the output parses. That is all. It says nothing about keys, types, or enums, so a missing field still gets through. In Chat Completions, your prompt must also contain the word "JSON", or the API returns an error.

Treat it as a fallback. If your model supports json_schema, OpenAI recommends that option instead. Schema-based openai structured output json mode gives you everything JSON mode does, plus enforcement of your exact shape.

Function calling is for actions

Function calling lets the model ask your code to run a tool. You describe the function with a parameter schema, and the model returns the arguments for that function. Add strict: true to the definition, and those arguments follow the same schema guarantee as a structured reply.

So the difference is intent, not strength. Use function calling when the model is choosing an action, such as looking up an order or issuing a refund. Use structured outputs when the reply itself is the deliverable, such as an extracted record or a label.

Use function calling when the model triggers an action, and structured outputs when the answer is the deliverable.

Side by side

Here is how the three compare on the two things that matter most. Check whether the output parses as JSON, and whether it matches your schema.

FeatureValid JSONMatches your schemaBest for
JSON modeYesNoOlder models, loose shapes
Function calling (strict)YesYesTool calls, agent actions
Structured Outputs (json_schema)YesYesFinal answers, extraction, classification

Agents often use two of these together. Let function calling drive the tool steps, then end the run with a structured reply for the user-facing answer. The Responses API accepts tools and a text.format schema in the same request, so you do not need separate calls.

Schema limits, refusals and common pitfalls

Strict mode accepts only a subset of JSON Schema, and it fails loudly when you step outside it. Knowing the edges before you design a schema saves you a round of 400 errors. With OpenAI structured output, these limits apply whether you use the Responses API or Chat Completions.

Hard limits to design around

OpenAI's Structured Outputs guide lists the caps below as of this writing. Check it again before you ship, because they change.

LimitCap
Root typeMust be an object, not anyOf
Total properties100 across the schema
Nesting depth5 levels
Enum values500 total
Names and enum text15,000 characters combined

Large extraction schemas hit these caps fast. When that happens, split one big call into two smaller ones or flatten nested objects. Also expect extra latency on the first request with a new schema, since the API processes it before generating. Later calls reuse that work.

Refusals are a valid response

Sometimes the model declines a request for safety reasons. In that case the API returns a refusal message instead of JSON, so the output will not match your schema. That is by design, and your code should treat it as a normal branch, not a crash.

Show the user a clear message, log the text, and move on. Do not retry the same prompt in a loop, because a refusal rarely changes on a second attempt.

A refusal is a valid response, so give it a code path before launch.

Pitfalls that burn teams

Most production bugs come from a short list of mistakes. Check your integration against these:

  • Parallel tool calls: strict function calling does not work with them. Set parallel_tool_calls to false.
  • Unsupported keywords: length, range, and pattern rules may be rejected or ignored, depending on the model. Enforce them in your own validator.
  • Key order: the model writes fields in schema order. Put a reasoning string before the answer field so the model thinks first, then commits.
  • Model support: older models such as gpt-4-0613 handle JSON mode only. Use gpt-4o-2024-08-06 or later for json_schema.
  • Streaming: partial JSON does not parse until it is complete. Use the SDK's stream helpers, or buffer the full text first.

None of these are hard to fix. They only hurt when you find them in production, so test one realistic request per schema before you deploy.

Putting structured outputs to work

OpenAI structured output turns "please reply in JSON" into a contract. Set strict: true and define the schema once in Pydantic, Zod, or raw JSON Schema. Your code can then parse every response without cleanup code and skip the format retry loop.

Remember what the contract leaves out. Check for refusals and truncation before you parse, and keep your own validation on top, because a schema guarantees shape, not truth. Respect the limits on depth, properties, and enums when you design.

Start small. Pick one extraction or classification task, write the schema, and test it on the exact model you plan to ship. The request format is portable, so you can run your schema on DeepSeek V4 Flash through Geodd's OpenAI-compatible API by changing only the base URL and model name.

The Chronicle / Bartosz Neuman
Keep reading

More from Geodd.

All articles