GLM 5.3

z-ai/glm-5.3-20260816

Z.ai: GLM 5.3 is a text-to-text model with a context length of 1,048,576 tokens. It supports tool use and reasoning. The model was released on August 14, 2026.

  • tool use
  • structured outputs
  • reasoning
  • open-weight

At a glance

providerz-ai
releasedAug 14, 2026
context1.3M
input / M$1.40

Specifications

What the API gives you

Context window
1.3M tokens
Max output
262K tokens
Input price
$1.40 / 1M tokens
Output price
$4.40 / 1M tokens
Cache read
$0.14 / 1M tokens
Modalities
text → text
Knowledge cutoff
Released
Aug 14, 2026

Measured evidence

Public benchmark record

Frontier-Bench
41.8%2026-08-14Claude Code, max effort
LiveBench Coding
79.0%2026-06-25
LiveBench Language
79.9%2026-06-25
LiveBench Math
87.9%2026-06-25
LiveBench Reasoning
85.8%2026-06-25
LMArena Creative Writing
1462 Elo2026-08-21
LMArena Creative Writing
1465 Elo2026-08-27
LMArena Creative Writing
1466 Elo2026-09-01
LMArena Long Query
1484 Elo2026-08-21
LMArena Long Query
1486 Elo2026-08-27
LMArena Long Query
1487 Elo2026-09-01
LMArena Math
1502 Elo2026-08-27
LMArena Math
1506 Elo2026-09-01
LMArena Text
1487 Elo2026-08-21
LMArena Text
1484 Elo2026-08-27
LMArena Text
1483 Elo2026-09-01
LMArena WebDev
1599 Elo2026-08-21
LMArena WebDev
1608 Elo2026-08-30
LMArena WebDev
1608 Elo2026-09-01
Terminal-Bench 2.1
41.8%2026-08-14max effort

LMArena data licensed CC-BY-4.0. Aider (Apache-2.0) and SWE-bench (MIT) are open. Design Arena uses blind human preference. Submitted scaffold plus model result. Compare the harness and effort settings, not only the model name.

Open-weight model on Hugging Face.

Task evidence

How broad is the support?

Each status is exact to this model version and the selected task. Official release claims remain provisional.

  • Best LLM for Competition MathEstablished evidence

    100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Business WritingEstablished evidence

    100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • 100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Marketing CopyEstablished evidence

    100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Chat & RoleplayEstablished evidence

    100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Creative WritingEstablished evidence

    100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for WritingEstablished evidence

    100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • 100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Math & ReasoningEstablished evidence

    100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • 100% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Code CompletionEstablished evidence

    81% of intended weighted task evidence across 2 independent sources. 4 of 4 configured evaluation feeds currently have results.

  • 67% of intended weighted task evidence across 2 independent sources. 3 of 4 configured evaluation feeds currently have results.

  • Best LLM for Fast SummariesEstablished evidence

    64% of intended weighted task evidence across 2 independent sources. 3 of 4 configured evaluation feeds currently have results.

  • Best Value LLMEstablished evidence

    64% of intended weighted task evidence across 2 independent sources. 3 of 4 configured evaluation feeds currently have results.

  • Best LLM for Autonomous AgentsEstablished evidence

    63% of intended weighted task evidence across 3 independent sources. 7 of 7 configured evaluation feeds currently have results.

  • 55% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from WMT translation. 4 of 5 configured evaluation feeds currently have results.

  • 55% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from WMT translation. 4 of 5 configured evaluation feeds currently have results.

  • Best LLM for TranslationProvisional evidence

    55% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from WMT translation. 4 of 5 configured evaluation feeds currently have results.

  • Best LLM for Agentic CodingProvisional evidence

    55% of intended weighted task evidence across 3 independent sources; broader independent confirmation is still missing from Design Arena Full-stack or Design Arena Web-apps. 8 of 8 configured evaluation feeds currently have results.

  • Best LLM for CodingProvisional evidence

    54% of intended weighted task evidence across 3 independent sources; broader independent confirmation is still missing from Aider Polyglot or Design Arena Coding. 8 of 8 configured evaluation feeds currently have results.

  • Best Cheap LLM for High VolumeProvisional evidence

    50% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Route reliability. 3 of 4 configured evaluation feeds currently have results.

  • Cheapest Capable LLMProvisional evidence

    50% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Route reliability. 3 of 4 configured evaluation feeds currently have results.

  • Best Local LLM for CodingProvisional evidence

    50% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Aider Polyglot or SWE-rebench. 5 of 5 configured evaluation feeds currently have results.

  • 40% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Aider Polyglot or SWE-bench Verified. 6 of 6 configured evaluation feeds currently have results.

  • 40% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from FACTS Grounding or LMArena Document. 4 of 6 configured evaluation feeds currently have results.

  • Best LLM for SummarizationProvisional evidence

    40% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from FACTS Grounding or LMArena Document. 4 of 6 configured evaluation feeds currently have results.

  • Best LLM for Frontend & UIProvisional evidence

    40% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Design Arena Coding or Design Arena UI. 5 of 5 configured evaluation feeds currently have results.

  • 39% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Berkeley Function Calling or Structured-output eval. 4 of 5 configured evaluation feeds currently have results.

  • 39% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Berkeley Function Calling or Structured-output eval. 4 of 5 configured evaluation feeds currently have results.

  • Best LLM for AI AgentsProvisional evidence

    33% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 7 of 7 configured evaluation feeds currently have results.

  • 31% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Route reliability or Structured-output eval. 2 of 4 configured evaluation feeds currently have results.

  • Best LLM for Massive ContextProvisional evidence

    25% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Document or LongBench v2. 3 of 4 configured evaluation feeds currently have results.

  • Best LLM for Hermes AgentProvisional evidence

    24% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for OpenHandsProvisional evidence

    24% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for GooseProvisional evidence

    24% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for OpenClawProvisional evidence

    24% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 5 of 5 configured evaluation feeds currently have results.

  • 20% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Vision or VLMEvalKit tasks. 3 of 4 configured evaluation feeds currently have results.

  • 18% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from FACTS Grounding or LMArena Document. 4 of 6 configured evaluation feeds currently have results.

  • 10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 4 of 5 configured evaluation feeds currently have results.

  • 10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 4 of 4 configured evaluation feeds currently have results.

  • Best LLM for OCR & DocumentsProvisional evidence

    10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Vision. 3 of 5 configured evaluation feeds currently have results.

  • 10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Vision. 3 of 5 configured evaluation feeds currently have results.

  • Best LLM for Grounded RAGProvisional evidence

    10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from CRAG or FACTS Grounding. 2 of 5 configured evaluation feeds currently have results.

  • Best LLM for Charts & DiagramsProvisional evidence

    10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Vision. 3 of 5 configured evaluation feeds currently have results.

  • Best LLM for Document ParsingProvisional evidence

    10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Vision. 3 of 5 configured evaluation feeds currently have results.

Recommendation coverage

Where this model appears

  1. #5Best LLM for Translation
  2. #7Best Open-Weight LLM for Self-Hosting

Catalog data via OpenRouter and models.dev. Last checked within the past 6 hours.