Recommendation for JSON & schema
JSON & Schema Output
The best LLM for JSON and schema output depends heavily on which models actually support structured output features, a capability some top performers inexplicably lack. For production data pipelines, the OpenAI GPT-5 models and Gemini 3 Flash Preview stand out as the safest choices because they still support structured outputs natively, while several top-ranked Claude models have dropped this feature entirely. This ranking prioritizes models that can reliably emit valid JSON conforming to a schema every time, not just general text quality. Several of the highest-ranking models on LMArena are surprisingly bad fits here because they no longer support structured outputs, making lower-ranked alternatives better for this specific use case.
About this recommendation
- Updated
- Jul 17, 2026
- Evidence through
- Jul 17, 2026
- Sources
- 9
- Revision
- v1
GPT-5.4 ranks #9 on LMArena and, critically, appears to still support structured outputs, making it the strongest overall choice for JSON generation in data pipelines. The combination of solid benchmark standing and retained structured output capability addresses the core need for schema-valid JSON that downstream systems can consume without parsing errors.
Best when: You need a high-quality general model that still supports structured outputs for reliable JSON schema conformance in production pipelines.
Tips
- Ranks #9 of 49 on LMArena overall, indicating strong general text quality (Elo 1470).
Watch out for
- Lacks explicit community testimony about its structured output performance specifically.
GPT-5.5 ties its sibling at #10 on LMArena and offers the same presumed structured output capability. It is functionally interchangeable with GPT-5.4 for JSON generation, with marginally lower benchmark standing being the only differentiator.
Best when: You want OpenAI's structured output feature and prefer the latest model, with Elo 1470 performance.
Tips
- Ranks #10 of 49 on LMArena overall (Elo 1470), competitive with GPT-5.4.
Watch out for
- No direct community evidence about its JSON or structured output behavior.
Gemini 3 Flash Preview ranks #12 on LMArena and has documented community testing for structured data extraction tasks. A user successfully tested it for extracting URLs and generating tables, providing concrete evidence it can produce structured output suitable for pipeline use.
Best when: You want a Gemini option with proven structured data extraction capabilities and community validation.
Tips
- Community tested for URL extraction and table generation, demonstrating real structured output use.
- Ranks #12 of 49 on LMArena overall (Elo 1466), respectable benchmark standing.
Watch out for
- Lower Elo (1466) than the top OpenAI and Anthropic models.
Claude Opus 4.6 holds the #1 spot on LMArena (Elo 1501) for general text quality, but its structured output support status is unclear. Unlike its newer siblings, there is no explicit mention of this model losing the feature, nor confirmation it retains it, introducing uncertainty for pipeline use.
Best when: You prioritize raw text quality and can tolerate uncertainty about structured output availability, or can implement workarounds.
Tips
- Top-ranked model overall at #1 of 49 on LMArena (Elo 1501).
Watch out for
- Status of structured output support is unconfirmed, creating risk for data pipeline use.
Claude Fable 5 ranks #2 overall on LMArena but is explicitly identified as having dropped structured output support. This makes it a poor fit for JSON schema output despite its excellent general capabilities, as users report frustration with the removal of a feature critical for passing LLM output to other APIs and libraries.
Best when: You need top-tier general text generation but not structured JSON output for downstream API integration.
Tips
- Second highest ranked model overall at #2 of 49 on LMArena (Elo 1493).
Watch out for
- Explicitly stopped supporting structured outputs, a feature users found incredibly useful for API integration.
- Community members frustrated by the loss of ability to reliably pass output to other libraries.
Claude Opus 4.7 ranks #3 overall but, like Fable 5, is confirmed to have dropped structured output support. Its strong Elo of 1490 does not compensate for the lack of a feature essential for reliable JSON generation in data pipelines.
Best when: General text tasks only. Avoid for structured output pipelines.
Tips
- Third highest ranked model overall at #3 of 49 on LMArena (Elo 1490).
Watch out for
- Explicitly stopped supporting structured outputs per community reports.
- Users report this removal makes it harder to pass LLM output directly to downstream APIs.
Claude Opus 4.8 rounds out the trio of Anthropic models that dropped structured output support. It ranks #13 overall (Elo 1462), lower than Fable and Opus 4.7, and shares the same limitation for JSON pipeline work.
Best when: General text generation where structured output is not required.
Tips
- Ranks #13 of 49 on LMArena overall (Elo 1462), decent benchmark standing.
Watch out for
- Named alongside Opus 4.7 and Fable 5 as having dropped structured output support.
- Users frustrated by inability to reliably pass output to other APIs and libraries.
Frequently asked
- Which Claude models support structured JSON output?
- Based on community reports, Claude Opus 4.7, Opus 4.8, and Fable 5 have all stopped supporting structured outputs, making them unreliable for programmatic data pipelines.
- Does model ranking on LMArena predict JSON output quality?
- Not necessarily. Several top-5 models on LMArena lack structured output support, so general text quality does not guarantee reliable schema-conforming JSON generation.
- What models work best for passing LLM output directly to other APIs?
- Models with native structured output support like OpenAI's GPT-5 series are ideal since they can emit JSON that validates against a schema, ready for downstream APIs and libraries.
- Is Gemini 3 Flash Preview good at structured data extraction?
- Community testing shows it can generate structured output formats like tables from extracted URLs, making it a viable option for data extraction tasks.
Sources
- 1
“Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 2
“Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 3
“Interesting! Where did you apply it? Can you show your output in more detail? It's more like a small script, and it's supposed to extract urls and generate a table. Here's my result in Claude Web for comparison: https: claude.ai public artifacts d76936f2-c97b-4bff-9205-2... Claude web finds a number of small discrepancies in the sources, which I manually crosschecked and seem consistent with a human mixing things up slightly. + I also tested in gemini 3 flash preview, which generates an actual…”
Kim_Bruning · Hacker News · May 15, 2026 - 4
“Ranks #12 of 49 on LMArena's overall text arena (Elo 1466), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 16, 2026 - 5
“Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 6
“Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 7
“I dont get why Opus 4.7, 4.8, and now Fable all stopped supporting structured outputs? Does no one else care about that? I find it incredibly useful to reliably pass LLM output directly to other APIs libraries”
coreylane · Hacker News · Jun 9, 2026 - 8
“Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 9
“Ranks #17 of 17 on LMArena's overall text arena (Elo 1472), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.