Recommendation for General assistant
a General Assistant
Our top recommendation for a General Assistant, based on the public evidence we track, is OpenAI: GPT-5.6 Sol.[1][2][3] Use for extended daily chat where you want the model to catch and correct its own mistakes mid-conversation. Watch out: Watch for transport errors and session termination in long conversations, with 19 transport-error events logged in one monitoring period including sessions with 311 messages. Anthropic: Claude Fable 5 is the next-ranked alternative. Use when you need strong adherence to custom guidelines without excessive length, as users report it avoids the 'diarrhea of the mouth' seen in Opus 5.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 27
- Revision
- v59
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
3 of 9
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LMArena Instruction Following | 30% | #11 | 19/21 |
| LiveBench Reasoning | 25% | #2 | 21/21 |
| LiveBench Instruction Following | 20% | #16 | 21/21 |
| LMArena Text | 15% | #10 | 19/21 |
| OpenRouter usage | 10% | 97/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic3 models
- Google2 models
- Meta1 model
- OpenAI1 model
- Qwen1 model
- xAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | GPT-5.6 SolOpenAI | 86 | 100% | 4 threads · 3 families · 2 cautions | #2 LiveBench Reasoning · #10 LMArena Text |
| 02 | Claude Fable 5Anthropic | 85 | 100% | 2 threads · 2 families · 1 cautions | #1 LMArena Text · #3 LMArena Instruction Following |
| 03 | Gemini 3.7 FlashGoogle | 85 | 100% | 1 threads · 1 families · 0 cautions | #2 LiveBench Instruction Following · #5 LMArena Instruction Following |
| 04 | Gemini 3.6 FlashGoogle | 83 | 100% | 6 threads · 3 families · 4 cautions | #7 LiveBench Instruction Following · #13 LMArena Text |
| 05 | Claude Opus 4.6Anthropic | 82 | 100% | 2 threads · 1 families · 2 cautions | #1 LMArena Instruction Following · #2 LMArena Text |
| 06 | Claude Opus 4.8Anthropic | 81 | 100% | 1 threads · 1 families · 0 cautions | #11 LiveBench Reasoning · #11 LMArena Instruction Following |
| 07 | Muse Spark 1.2Meta | 80 | 100% | 1 threads · 1 families · 1 cautions | #4 LMArena Text · #7 LiveBench Reasoning |
| 08 | Grok 4.6xAI | 78 | 100% | 2 threads · 2 families · 1 cautions | #6 LiveBench Reasoning · #15 LiveBench Instruction Following |
| 09 | Qwen3.8 27BQwen | 76 | 100% | 2 threads · 2 families · 0 cautions | #13 LiveBench Instruction Following · #30 LMArena Instruction Following |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
GPT-5.6 Sol shows reliable self-correction and adaptive tone in daily conversation, though transport errors and UI friction affect the experience.
Best when: Use for extended daily chat where you want the model to catch and correct its own mistakes mid-conversation.
Tips
- Use for extended daily chat where you want the model to catch and correct its own mistakes mid-conversation.
- Use with 'medium thinking' enabled when you need a balance between response speed and quality on complex questions.
Watch out for
- Watch for transport errors and session termination in long conversations, with 19 transport-error events logged in one monitoring period including sessions with 311 messages.
- Check that you're actually on Sol and not 5.5 Instant, as the effort selector may default to the cheaper model without clear indication.
Claude Fable 5 ranks highly for instruction following with controlled verbosity compared to its Opus sibling, though safeguard triggers can downgrade sessions unexpectedly.
Best when: Use when you need strong adherence to custom guidelines without excessive length, as users report it avoids the 'diarrhea of the mouth' seen in Opus 5.
Tips
- Use when you need strong adherence to custom guidelines without excessive length, as users report it avoids the 'diarrhea of the mouth' seen in Opus 5.
- Use for instruction-following tasks like paraphrasing and summarization where LiveBench scores 75.77% (#5 of 51).
Watch out for
- Watch for false positive safety flags that can silently downgrade your session to Opus 4 without clear notification.
Gemini 3.7 Flash achieves the #2 LiveBench Instruction Following score with strong human preference rankings, though API integration quirks persist.
Best when: Use for instruction-following tasks where benchmark standing matters, as it scores 79.93% on LiveBench Instruction Following (#2 of 51).
Tips
- Use for instruction-following tasks where benchmark standing matters, as it scores 79.93% on LiveBench Instruction Following (#2 of 51).
- Use when you want a model that ranks #5 in human preference for instruction following on LMArena.
Watch out for
- Watch for model list formatting issues when integrating via OpenAI-compatible endpoints, as the API may return 'models/' prefixes that cause 404 errors.
Gemini 3.6 Flash offers solid instruction-following performance but carries strict rate limits and streaming quirks that constrain real-world usage.
Best when: Use for code generation tasks where empty replies have not been observed, unlike some other Gemini variants at small max_tokens settings.
Tips
- Use for code generation tasks where empty replies have not been observed, unlike some other Gemini variants at small max_tokens settings.
Watch out for
- Avoid for high-volume applications on free tier, as the 20 requests per day per project limit translates to roughly 10 actual messages when using multi-call workflows.
- Watch for missing completion_tokens in usage objects when streaming heavy-reasoning responses, which can crash strict OpenAI-compatible clients.
- Watch for 'No response: RAW' errors in book chat features when using via OpenRouter.
Claude Opus 4.6 tops LMArena's instruction-following category but fabricates explanations about its own reasoning when challenged.
Best when: Use when you need the highest human-rated instruction following (#1 on LMArena with Elo 1523) for tasks like complex formatting or multi-step directions.
Tips
- Use when you need the highest human-rated instruction following (#1 on LMArena with Elo 1523) for tasks like complex formatting or multi-step directions.
- Use for reasoning tasks where it scores 88.67% on LiveBench Reasoning.
Watch out for
- Watch for fabricated justifications when you question its outputs, as it may generate false explanations about its own reasoning process rather than admit uncertainty.
Claude Opus 4.8 shows mid-tier instruction-following scores with deployment via multiple cloud routes.
Best when: Use when you need Anthropic's SDK directly rather than via OpenRouter, as some implementations specifically target this model endpoint.
Tips
- Use when you need Anthropic's SDK directly rather than via OpenRouter, as some implementations specifically target this model endpoint.
Watch out for
- Note the significant drop in LiveBench Instruction Following to 72.03% (#14 of 51) compared to other Claude variants.
Muse Spark 1.2 combines strong reasoning scores with limited deployment routes and reported availability issues.
Best when: Use for reasoning tasks where it scores 90% on LiveBench Reasoning (#7 of 51).
Tips
- Use for reasoning tasks where it scores 90% on LiveBench Reasoning (#7 of 51).
Watch out for
- Watch for availability issues on free tiers, with reports of 181001 errors and HTTP 400 BadRequestError indicating model unavailability.
Grok 4.6 shows strong reasoning performance but suffers from integration bugs and conversation compaction failures.
Best when: Use for reasoning tasks where it scores 90.51% on LiveBench Reasoning (#6 of 51).
Tips
- Use for reasoning tasks where it scores 90.51% on LiveBench Reasoning (#6 of 51).
Watch out for
- Watch for IndexError crashes when streaming redacted reasoning blocks via Bedrock ConverseStream, as the SDK lacks handling for encrypted reasoning content.
- Watch for 'Could not compact conversation' errors that terminate sessions without output, particularly in longer conversations.
Qwen3.8 27B serves as a capable local open-weight option with strong context window support, though instruction-following benchmarks trail cloud alternatives.
Best when: Use for local deployment with 262K context window when you need resident low-TTFT interactive traffic without cloud dependency.
Tips
- Use for local deployment with 262K context window when you need resident low-TTFT interactive traffic without cloud dependency.
Watch out for
- Note the lower LiveBench Instruction Following score of 72.66% (#13 of 51) compared to top cloud alternatives.
Frequently asked
- What is the top-ranked model for a General Assistant?
- OpenAI: GPT-5.6 Sol ranks first in the current evidence-weighted comparison. Use for extended daily chat where you want the model to catch and correct its own mistakes mid-conversation.[1]
- What should I watch out for with OpenAI: GPT-5.6 Sol?
- Watch for transport errors and session termination in long conversations, with 19 transport-error events logged in one monitoring period including sessions with 311 messages.[2]
- What is an alternative to OpenAI: GPT-5.6 Sol?
- Anthropic: Claude Fable 5 is the next-ranked option. Use when you need strong adherence to custom guidelines without excessive length, as users report it avoids the 'diarrhea of the mouth' seen in Opus 5.[3]
Sources
- 1
“5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.”
leokennis · Hacker News · Aug 18, 2026 - 2
“## 背景 Pi Web Desktop 中 openai-codex(ChatGPT 订阅 OAuth)会话频繁出现「输出一半停止 / fetch failed / terminated」,其他模型(aliyun/deepseek)偶发但次数少。 ## 实锤数据(2026-08-01,scripts/session-stops.mjs 修复后全量扫描 185 会话) - transport-error 共 37 条:openai-codex/GPT 19(gpt-5.6-sol×11 + gpt-5.6-terra×8)、aliyun 9、pi-router 8、deepseek 1。 - 最大异常会话:311 消息 / ~2.2M tokens,会话开头连续 6 次 fetch failed/terminated。 - compaction 仅 1 次且在非中断点,**不是主因**;主因是 transport error。 ## 工具修复(commit bfe020d) session-stops.mjs 原判定只在独立 type=error 事件时识别 transport-er…”
dust617 · GitHub · Aug 1, 2026 - 3
“Which Claude 5? Opus 5 does seem to have diarrhea of the mouth. But Fable 5 hasn't been so bad for me. Or perhaps it is just better at adhering to my guidelines.”
rootusrootus · Hacker News · Aug 20, 2026 - 4
“For me on a paid plan, the effort indicator was hidden and the model was 5.5 instant. I had to press the + button to select “think harder” before the dial that allowed Sol medium or high to be selected to show. It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?”
aryehof · Hacker News · Aug 7, 2026 - 5
“Scores 75.77% on LiveBench Instruction Following (#5 of 51), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 6
“**Bug Description** you flagged something completely innocent and dropped back to opus. it seems pointless to even have fable /model claude-fable-5 **Environment Info** - Platform: darwin - Terminal: ghostty - Version: 2.1.220 - Feedback ID: 5aaf422a-7737-4231-94c4-737c4df0264b **Errors**”
elyochola · GitHub · Aug 2, 2026 - 7
“Scores 79.93% on LiveBench Instruction Following (#2 of 51), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 8
“Ranks #5 of 144 on LMArena's instruction-following category (Elo 1485), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 9
“### Hermes Web UI Version v0.6043 ### Hermes Agent Version v0.20.3 ### Bug Description 在cli中添加的Google AI Studio模型在web-ui中看不到,如果使用web-ui内置的Google AI Studio添加(base为https://generativelanguage.googleapis.com/v1beta/openai),获取到的模型列表全部带“models/”前缀,如“models/gemini-3.7-flash”,导致对话时gemini返回404 如果手动添加未列出的模型名为“gemini-3.7-flash”则可以正常对话 ### Steps to Reproduce 1、web-ui中“模型”设置页面点击右上角“添加 Provider” 2、“选择 Provider”中选择“Google AI Studio” 3、填入API Key,“默认模型”处点击“获取” 4、点击“添加” 5、新建对话,选择该模型列表中任一模型,输入对话,对话框返回“Error: Gemi…”
hspmanbu · GitHub · Aug 18, 2026 - 10
“Model confirmed by the reporter as **`gemini-3.6-flash`**, not `gemini-3.5-flash`. Acceptance updated. Worth noting which way that cuts: `gemini-3.6-flash` is the most expensive of the three configured models and the one #32 is filed against. It handled editor chat code generation with no empty replies — consistent with #32, which only bites at small `max_tokens`, and Continue's chat role does not set one.”
ujjawalmisra · GitHub · Aug 5, 2026 - 11
“## 문제 요약 `gemini-3.6-flash` 무료 티어가 프로젝트+모델 기준 하루 20건 요청으로 제한되어 있는데(`GenerateRequestsPerDayPerProjectPerModel-FreeTier`), 메시지 하나당 LLM을 최소 2번(1차 긴급도 판단 + RAG 재검색 후 2차 최종 판단) 호출하므로 실질적으로 하루 메시지 10개 정도면 한도가 참. 실사용 메신저 봇으로는 감당 불가. ## 에러 로그 ## 해결 방안 (적용함) 채팅 생성 LLM을 Gemini에서 **Groq**(`llama-3.3-70b-versatile`)로 전환. Groq 무료 티어는 분당 요청수 기준이라 하루 단위로 막히는 지금 상황보다 훨씬 넉넉함. - `app/pipelines/shared/llm.py` 신규 — `get_chat_llm()` 공용 팩토리 - `realtime_graph.py`(3곳), `feedback_graph.py`(2곳), `memory_gc_graph.py`(1곳)…”
oh0227 · GitHub · Jul 24, 2026 - 12
“## Problem When proxying Antigravity/Gemini responses to the OpenAI chat-completions format, the `usage` object in a streamed chunk (and potentially non-stream responses) can be emitted **without** `completion_tokens`. Strict OpenAI clients that deserialize `usage` with required fields (e.g. Rust serde in Grok CLI) fail the whole turn: Observed with `gemini-3.6-flash-high` via Antigravity OAuth. Heavy-reasoning responses appear to trigger it: upstream sends a `usageMetadata` chunk containing `p…”
amikai · GitHub · Jul 21, 2026 - 13
“**Describe the bug** Returns “No response: RAW” whenever using ‘book chat’ feature. First it does show a streaming response to my prompt, but when it finishes I get the error. **To Reproduce** Using Openrouter and Gemini 3.6 Flash, when I use the ‘Book Chat’. **Expected behavior** It to show me the response to the chat I submitted. **e-reader (please complete the following information):** - Device: Musnap Ocean - OS: Android 14 - KOReader **Additional context** Haven’t had any issues with any o…”
rockettemorton · GitHub · Jul 24, 2026 - 14
“Ranks #1 of 144 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 15
“Scores 88.67% on LiveBench Reasoning (#14 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 16
“### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…”
NStestUser1954 · GitHub · Jul 22, 2026 - 17
“### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…”
NStestUser1954 · GitHub · Jul 22, 2026 - 18
“## Contexto Existem duas implementações independentes do mesmo endpoint de chat ("Consciência Assistida" do AI Lab): - `server.js` (rota `/api/ai-lab-chat`, usada apenas em dev via `npm run dev` — Express local, provedor **OpenRouter**, `OPENROUTER_MODEL` default `openai/gpt-4o-mini`) - `api/ai-lab-chat.js` (Vercel serverless function, provedor **Anthropic** via `@anthropic-ai/sdk`, modelo `claude-opus-4-8`) ## Evidência `server.js`: `api/ai-lab-chat.js`: O system prompt já **divergiu de fato**…”
luksjfernandes-ctrl · GitHub · Jul 26, 2026 - 19
“Scores 72.03% on LiveBench Instruction Following (#14 of 51), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 20
“Scores 90% on LiveBench Reasoning (#7 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 21
“### 分支选择 dev-stable ### 模块选择 LLM wiki ### Checklist - [x] 我已经搜索过相关问题,但没有得到预期的帮助。 - [x] 最新版本中该错误尚未修复。 - [x] 请注意,如果您提交的Bug描述缺少相应的环境信息和最小可复现的demo,我们将很难复现和解决该问题,从而降低收到反馈的可能性,甚至该问题将被关闭。 ### 🐞 问题详细描述 ●内置模型来源:后端 Zen 提供的 9 个 free 模型(前端"内置免费模型"列表) 问题汇总 # 前端显示名 后端标识符(已知/待确认) 现象 错误码 错误类型 1 DeepSeek V4 Flash deepseek-v4-flash-free 调用报错 181001 / HTTP 400 BadRequestError: Model is unavailable 2 OX Alpha Free 待确认(推测带-free后缀) 对话无反应 / 无输出 无(无报错) 静默失败(疑似超时或未捕获错误) 3 Muse Spark 1.2 待确认(推测带-free后缀) 调用报错 181001 / H…”
openjiuwen-collaboration-bot[bot] · GitHub · Aug 24, 2026 - 22
“Scores 90.51% on LiveBench Reasoning (#6 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 23
“## Description `ChatBedrockConverse.stream()` crashes with `IndexError: list index out of range` when the model streams redacted reasoning blocks. Amazon Bedrock's `ConverseStream` can deliver reasoning as either `reasoningContent.text` (plain) or `reasoningContent.redactedContent` (encrypted bytes). `_bedrock_to_lc()` handles the first form but has no branch for the second, so it returns an empty list. `_parse_stream_event()` immediately indexes `[0]` into that result and raises. xAI Grok 4.6…”
Great-Root · GitHub · Aug 21, 2026 - 24
“添加kimi-k3和grok-4.6后,在使用时会报Could not compact conversation然后终止,没有任何输出。而使用kimi-k2.7-highspeed和glm-5.3则没有这样的问题。请问是不是我配置的有问题呢?  ”
hhr114 · GitHub · Aug 21, 2026 - 25
“## Destination A single lean llama.cpp `llama-server` (HIP, gfx1100) serving Qwen3.8-27B (unsloth `UD-Q3_K_XL`) on the RX 7900 GRE — KV-cache pre-sized to fit in VRAM so it physically cannot OOM, exposed OpenAI-compatible, fronted by the Hollama browser UI as the opencode replacement, and launched under a systemd unit with a hard memory cap. "Done" = a usable model I can chat/code with: no swap, no 27GB RAM spike, correct context actually provisioned. ## Notes - Domain: local LLM serving, llama…”
darylerivers · GitHub · Aug 19, 2026 - 26
“## Why Two workload shapes with opposite requirements are currently served by one always-on GPU: - **Sprinkled interactive traffic** (Discord, chat, vision, classifier, pi turns). Needs a resident model and low TTFT. Served well today by the 4090 running Qwen3.8-27B at 262K context. - **Clumped batch work** (`model-bench` sweeps, bulk classification, embeddings backfills, autonomous queue jobs). Latency-insensitive, preemption-tolerant, and wants a much larger model than the 4090 can hold. The…”
jomcgi · GitHub · Aug 26, 2026 - 27
“Scores 72.66% on LiveBench Instruction Following (#13 of 51), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.