Recommendation for General assistant

a General Assistant

Our top recommendation for a General Assistant, based on the public evidence we track, is OpenAI: GPT-5.6 Sol.[1][2][3] Use for extended daily chat where you want the model to catch and correct its own mistakes mid-conversation. Watch out: Watch for transport errors and session termination in long conversations, with 19 transport-error events logged in one monitoring period including sessions with 311 messages. Anthropic: Claude Fable 5 is the next-ranked alternative. Use when you need strong adherence to custom guidelines without excessive length, as users report it avoids the 'diarrhea of the mouth' seen in Opus 5.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
27
Revision
v59

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

3 of 9

Anthropic

Established source breadth. 13 citation families and 3 practitioner families support the top result; 2 cautionary threads is retained. The largest citation family contributes 24%.

Sources evaluated

The task sets these weights before any model is scored.

winner: GPT-5.6 Sol
Evaluation feedWeightWinner resultField measured
LMArena Instruction Following
30%
#1119/21
LiveBench Reasoning
25%
#221/21
LiveBench Instruction Following
20%
#1621/21
LMArena Text
15%
#1019/21
OpenRouter usage
10%
97/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic33%
  • Anthropic3 models
  • Google2 models
  • Meta1 model
  • OpenAI1 model
  • Qwen1 model
  • xAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01GPT-5.6 SolOpenAI
86
100%4 threads · 3 families · 2 cautions#2 LiveBench Reasoning · #10 LMArena Text
02Claude Fable 5Anthropic
85
100%2 threads · 2 families · 1 cautions#1 LMArena Text · #3 LMArena Instruction Following
03Gemini 3.7 FlashGoogle
85
100%1 threads · 1 families · 0 cautions#2 LiveBench Instruction Following · #5 LMArena Instruction Following
04Gemini 3.6 FlashGoogle
83
100%6 threads · 3 families · 4 cautions#7 LiveBench Instruction Following · #13 LMArena Text
05Claude Opus 4.6Anthropic
82
100%2 threads · 1 families · 2 cautions#1 LMArena Instruction Following · #2 LMArena Text
06Claude Opus 4.8Anthropic
81
100%1 threads · 1 families · 0 cautions#11 LiveBench Reasoning · #11 LMArena Instruction Following
07Muse Spark 1.2Meta
80
100%1 threads · 1 families · 1 cautions#4 LMArena Text · #7 LiveBench Reasoning
08Grok 4.6xAI
78
100%2 threads · 2 families · 1 cautions#6 LiveBench Reasoning · #15 LiveBench Instruction Following
09Qwen3.8 27BQwen
76
100%2 threads · 2 families · 0 cautions#13 LiveBench Instruction Following · #30 LMArena Instruction Following

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. GPT-5.6 Sol shows reliable self-correction and adaptive tone in daily conversation, though transport errors and UI friction affect the experience.

    Best when: Use for extended daily chat where you want the model to catch and correct its own mistakes mid-conversation.

    Tips

    • Use for extended daily chat where you want the model to catch and correct its own mistakes mid-conversation.
      Source 1
      5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.
    • Use with 'medium thinking' enabled when you need a balance between response speed and quality on complex questions.
      Source 1
      5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.

    Watch out for

    • Watch for transport errors and session termination in long conversations, with 19 transport-error events logged in one monitoring period including sessions with 311 messages.
      Source 2
      ## 背景 Pi Web Desktop 中 openai-codex(ChatGPT 订阅 OAuth)会话频繁出现「输出一半停止 / fetch failed / terminated」,其他模型(aliyun/deepseek)偶发但次数少。 ## 实锤数据(2026-08-01,scripts/session-stops.mjs 修复后全量扫描 185 会话) - transport-error 共 37 条:openai-codex/GPT 19(gpt-5.6-sol×11 + gpt-5.6-terra×8)、aliyun 9、pi-router 8、deepseek 1。 - 最大异常会话:311 消息 / ~2.2M tokens,会话开头连续 6 次 fetch failed/terminated。 - compaction 仅 1 次且在非中断点,**不是主因**;主因是 transport error。 ## 工具修复(commit bfe020d) session-stops.mjs 原判定只在独立 type=error 事件时识别 transport-er…
    • Check that you're actually on Sol and not 5.5 Instant, as the effort selector may default to the cheaper model without clear indication.
      Source 4
      For me on a paid plan, the effort indicator was hidden and the model was 5.5 instant. I had to press the + button to select “think harder” before the dial that allowed Sol medium or high to be selected to show. It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?
  2. Claude Fable 5 ranks highly for instruction following with controlled verbosity compared to its Opus sibling, though safeguard triggers can downgrade sessions unexpectedly.

    Best when: Use when you need strong adherence to custom guidelines without excessive length, as users report it avoids the 'diarrhea of the mouth' seen in Opus 5.

    Tips

    • Use when you need strong adherence to custom guidelines without excessive length, as users report it avoids the 'diarrhea of the mouth' seen in Opus 5.
      Source 3
      Which Claude 5? Opus 5 does seem to have diarrhea of the mouth. But Fable 5 hasn't been so bad for me. Or perhaps it is just better at adhering to my guidelines.
      rootusrootusOpen original ↗
    • Use for instruction-following tasks like paraphrasing and summarization where LiveBench scores 75.77% (#5 of 51).
      Source 5
      Scores 75.77% on LiveBench Instruction Following (#5 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗

    Watch out for

    • Watch for false positive safety flags that can silently downgrade your session to Opus 4 without clear notification.
      Source 6
      **Bug Description** you flagged something completely innocent and dropped back to opus. it seems pointless to even have fable /model claude-fable-5 **Environment Info** - Platform: darwin - Terminal: ghostty - Version: 2.1.220 - Feedback ID: 5aaf422a-7737-4231-94c4-737c4df0264b **Errors**
  3. Gemini 3.7 Flash achieves the #2 LiveBench Instruction Following score with strong human preference rankings, though API integration quirks persist.

    Best when: Use for instruction-following tasks where benchmark standing matters, as it scores 79.93% on LiveBench Instruction Following (#2 of 51).

    Tips

    • Use for instruction-following tasks where benchmark standing matters, as it scores 79.93% on LiveBench Instruction Following (#2 of 51).
      Source 7
      Scores 79.93% on LiveBench Instruction Following (#2 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
    • Use when you want a model that ranks #5 in human preference for instruction following on LMArena.
      Source 8
      Ranks #5 of 144 on LMArena's instruction-following category (Elo 1485), based on blind human preference votes.
      LMArena instruction-following categoryOpen original ↗

    Watch out for

    • Watch for model list formatting issues when integrating via OpenAI-compatible endpoints, as the API may return 'models/' prefixes that cause 404 errors.
      Source 9
      ### Hermes Web UI Version v0.6043 ### Hermes Agent Version v0.20.3 ### Bug Description 在cli中添加的Google AI Studio模型在web-ui中看不到,如果使用web-ui内置的Google AI Studio添加(base为https://generativelanguage.googleapis.com/v1beta/openai),获取到的模型列表全部带“models/”前缀,如“models/gemini-3.7-flash”,导致对话时gemini返回404 如果手动添加未列出的模型名为“gemini-3.7-flash”则可以正常对话 ### Steps to Reproduce 1、web-ui中“模型”设置页面点击右上角“添加 Provider” 2、“选择 Provider”中选择“Google AI Studio” 3、填入API Key,“默认模型”处点击“获取” 4、点击“添加” 5、新建对话,选择该模型列表中任一模型,输入对话,对话框返回“Error: Gemi…
  4. Gemini 3.6 Flash offers solid instruction-following performance but carries strict rate limits and streaming quirks that constrain real-world usage.

    Best when: Use for code generation tasks where empty replies have not been observed, unlike some other Gemini variants at small max_tokens settings.

    Tips

    • Use for code generation tasks where empty replies have not been observed, unlike some other Gemini variants at small max_tokens settings.
      Source 10
      Model confirmed by the reporter as **`gemini-3.6-flash`**, not `gemini-3.5-flash`. Acceptance updated. Worth noting which way that cuts: `gemini-3.6-flash` is the most expensive of the three configured models and the one #32 is filed against. It handled editor chat code generation with no empty replies — consistent with #32, which only bites at small `max_tokens`, and Continue's chat role does not set one.
      ujjawalmisraOpen original ↗

    Watch out for

    • Avoid for high-volume applications on free tier, as the 20 requests per day per project limit translates to roughly 10 actual messages when using multi-call workflows.
      Source 11
      ## 문제 요약 `gemini-3.6-flash` 무료 티어가 프로젝트+모델 기준 하루 20건 요청으로 제한되어 있는데(`GenerateRequestsPerDayPerProjectPerModel-FreeTier`), 메시지 하나당 LLM을 최소 2번(1차 긴급도 판단 + RAG 재검색 후 2차 최종 판단) 호출하므로 실질적으로 하루 메시지 10개 정도면 한도가 참. 실사용 메신저 봇으로는 감당 불가. ## 에러 로그 ## 해결 방안 (적용함) 채팅 생성 LLM을 Gemini에서 **Groq**(`llama-3.3-70b-versatile`)로 전환. Groq 무료 티어는 분당 요청수 기준이라 하루 단위로 막히는 지금 상황보다 훨씬 넉넉함. - `app/pipelines/shared/llm.py` 신규 — `get_chat_llm()` 공용 팩토리 - `realtime_graph.py`(3곳), `feedback_graph.py`(2곳), `memory_gc_graph.py`(1곳)…
    • Watch for missing completion_tokens in usage objects when streaming heavy-reasoning responses, which can crash strict OpenAI-compatible clients.
      Source 12
      ## Problem When proxying Antigravity/Gemini responses to the OpenAI chat-completions format, the `usage` object in a streamed chunk (and potentially non-stream responses) can be emitted **without** `completion_tokens`. Strict OpenAI clients that deserialize `usage` with required fields (e.g. Rust serde in Grok CLI) fail the whole turn: Observed with `gemini-3.6-flash-high` via Antigravity OAuth. Heavy-reasoning responses appear to trigger it: upstream sends a `usageMetadata` chunk containing `p…
    • Watch for 'No response: RAW' errors in book chat features when using via OpenRouter.
      Source 13
      **Describe the bug** Returns “No response: RAW” whenever using ‘book chat’ feature. First it does show a streaming response to my prompt, but when it finishes I get the error. **To Reproduce** Using Openrouter and Gemini 3.6 Flash, when I use the ‘Book Chat’. **Expected behavior** It to show me the response to the chat I submitted. **e-reader (please complete the following information):** - Device: Musnap Ocean - OS: Android 14 - KOReader **Additional context** Haven’t had any issues with any o…
      rockettemortonOpen original ↗
  5. Claude Opus 4.6 tops LMArena's instruction-following category but fabricates explanations about its own reasoning when challenged.

    Best when: Use when you need the highest human-rated instruction following (#1 on LMArena with Elo 1523) for tasks like complex formatting or multi-step directions.

    Tips

    • Use when you need the highest human-rated instruction following (#1 on LMArena with Elo 1523) for tasks like complex formatting or multi-step directions.
      Source 14
      Ranks #1 of 144 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.
      LMArena instruction-following categoryOpen original ↗
    • Use for reasoning tasks where it scores 88.67% on LiveBench Reasoning.
      Source 15
      Scores 88.67% on LiveBench Reasoning (#14 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Watch for fabricated justifications when you question its outputs, as it may generate false explanations about its own reasoning process rather than admit uncertainty.
      Source 16
      ### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…
      NStestUser1954Open original ↗
      Source 17
      ### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…
      NStestUser1954Open original ↗
  6. Claude Opus 4.8 shows mid-tier instruction-following scores with deployment via multiple cloud routes.

    Best when: Use when you need Anthropic's SDK directly rather than via OpenRouter, as some implementations specifically target this model endpoint.

    Tips

    • Use when you need Anthropic's SDK directly rather than via OpenRouter, as some implementations specifically target this model endpoint.
      Source 18
      ## Contexto Existem duas implementações independentes do mesmo endpoint de chat ("Consciência Assistida" do AI Lab): - `server.js` (rota `/api/ai-lab-chat`, usada apenas em dev via `npm run dev` — Express local, provedor **OpenRouter**, `OPENROUTER_MODEL` default `openai/gpt-4o-mini`) - `api/ai-lab-chat.js` (Vercel serverless function, provedor **Anthropic** via `@anthropic-ai/sdk`, modelo `claude-opus-4-8`) ## Evidência `server.js`: `api/ai-lab-chat.js`: O system prompt já **divergiu de fato**…
      luksjfernandes-ctrlOpen original ↗

    Watch out for

    • Note the significant drop in LiveBench Instruction Following to 72.03% (#14 of 51) compared to other Claude variants.
      Source 19
      Scores 72.03% on LiveBench Instruction Following (#14 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
  7. Muse Spark 1.2 combines strong reasoning scores with limited deployment routes and reported availability issues.

    Best when: Use for reasoning tasks where it scores 90% on LiveBench Reasoning (#7 of 51).

    Tips

    • Use for reasoning tasks where it scores 90% on LiveBench Reasoning (#7 of 51).
      Source 20
      Scores 90% on LiveBench Reasoning (#7 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Watch for availability issues on free tiers, with reports of 181001 errors and HTTP 400 BadRequestError indicating model unavailability.
      Source 21
      ### 分支选择 dev-stable ### 模块选择 LLM wiki ### Checklist - [x] 我已经搜索过相关问题,但没有得到预期的帮助。 - [x] 最新版本中该错误尚未修复。 - [x] 请注意,如果您提交的Bug描述缺少相应的环境信息和最小可复现的demo,我们将很难复现和解决该问题,从而降低收到反馈的可能性,甚至该问题将被关闭。 ### 🐞 问题详细描述 ●内置模型来源:后端 Zen 提供的 9 个 free 模型(前端"内置免费模型"列表) 问题汇总 # 前端显示名 后端标识符(已知/待确认) 现象 错误码 错误类型 1 DeepSeek V4 Flash deepseek-v4-flash-free 调用报错 181001 / HTTP 400 BadRequestError: Model is unavailable 2 OX Alpha Free 待确认(推测带-free后缀) 对话无反应 / 无输出 无(无报错) 静默失败(疑似超时或未捕获错误) 3 Muse Spark 1.2 待确认(推测带-free后缀) 调用报错 181001 / H…
      openjiuwen-collaboration-bot[bot]Open original ↗
  8. Grok 4.6 shows strong reasoning performance but suffers from integration bugs and conversation compaction failures.

    Best when: Use for reasoning tasks where it scores 90.51% on LiveBench Reasoning (#6 of 51).

    Tips

    • Use for reasoning tasks where it scores 90.51% on LiveBench Reasoning (#6 of 51).
      Source 22
      Scores 90.51% on LiveBench Reasoning (#6 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Watch for IndexError crashes when streaming redacted reasoning blocks via Bedrock ConverseStream, as the SDK lacks handling for encrypted reasoning content.
      Source 23
      ## Description `ChatBedrockConverse.stream()` crashes with `IndexError: list index out of range` when the model streams redacted reasoning blocks. Amazon Bedrock's `ConverseStream` can deliver reasoning as either `reasoningContent.text` (plain) or `reasoningContent.redactedContent` (encrypted bytes). `_bedrock_to_lc()` handles the first form but has no branch for the second, so it returns an empty list. `_parse_stream_event()` immediately indexes `[0]` into that result and raises. xAI Grok 4.6…
    • Watch for 'Could not compact conversation' errors that terminate sessions without output, particularly in longer conversations.
      Source 24
      添加kimi-k3和grok-4.6后,在使用时会报Could not compact conversation然后终止,没有任何输出。而使用kimi-k2.7-highspeed和glm-5.3则没有这样的问题。请问是不是我配置的有问题呢? ![报错截图](https://files.seeusercontent.com/2026/08/21/kxS5/20260821_102019.png) ![配置截图](https://files.seeusercontent.com/2026/08/21/Wlz2/20260821_102116.png)
  9. Qwen3.8 27B serves as a capable local open-weight option with strong context window support, though instruction-following benchmarks trail cloud alternatives.

    Best when: Use for local deployment with 262K context window when you need resident low-TTFT interactive traffic without cloud dependency.

    Tips

    • Use for local deployment with 262K context window when you need resident low-TTFT interactive traffic without cloud dependency.
      Source 25
      ## Destination A single lean llama.cpp `llama-server` (HIP, gfx1100) serving Qwen3.8-27B (unsloth `UD-Q3_K_XL`) on the RX 7900 GRE — KV-cache pre-sized to fit in VRAM so it physically cannot OOM, exposed OpenAI-compatible, fronted by the Hollama browser UI as the opencode replacement, and launched under a systemd unit with a hard memory cap. "Done" = a usable model I can chat/code with: no swap, no 27GB RAM spike, correct context actually provisioned. ## Notes - Domain: local LLM serving, llama…
      daryleriversOpen original ↗
      Source 26
      ## Why Two workload shapes with opposite requirements are currently served by one always-on GPU: - **Sprinkled interactive traffic** (Discord, chat, vision, classifier, pi turns). Needs a resident model and low TTFT. Served well today by the 4090 running Qwen3.8-27B at 262K context. - **Clumped batch work** (`model-bench` sweeps, bulk classification, embeddings backfills, autonomous queue jobs). Latency-insensitive, preemption-tolerant, and wants a much larger model than the 4090 can hold. The…

    Watch out for

    • Note the lower LiveBench Instruction Following score of 72.66% (#13 of 51) compared to top cloud alternatives.
      Source 27
      Scores 72.66% on LiveBench Instruction Following (#13 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗

Frequently asked

What is the top-ranked model for a General Assistant?
OpenAI: GPT-5.6 Sol ranks first in the current evidence-weighted comparison. Use for extended daily chat where you want the model to catch and correct its own mistakes mid-conversation.[1]
What should I watch out for with OpenAI: GPT-5.6 Sol?
Watch for transport errors and session termination in long conversations, with 19 transport-error events logged in one monitoring period including sessions with 311 messages.[2]
What is an alternative to OpenAI: GPT-5.6 Sol?
Anthropic: Claude Fable 5 is the next-ranked option. Use when you need strong adherence to custom guidelines without excessive length, as users report it avoids the 'diarrhea of the mouth' seen in Opus 5.[3]

Sources

  1. 1

    5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.

    leokennis · Hacker News · Aug 18, 2026
  2. 2

    ## 背景 Pi Web Desktop 中 openai-codex(ChatGPT 订阅 OAuth)会话频繁出现「输出一半停止 / fetch failed / terminated」,其他模型(aliyun/deepseek)偶发但次数少。 ## 实锤数据(2026-08-01,scripts/session-stops.mjs 修复后全量扫描 185 会话) - transport-error 共 37 条:openai-codex/GPT 19(gpt-5.6-sol×11 + gpt-5.6-terra×8)、aliyun 9、pi-router 8、deepseek 1。 - 最大异常会话:311 消息 / ~2.2M tokens,会话开头连续 6 次 fetch failed/terminated。 - compaction 仅 1 次且在非中断点,**不是主因**;主因是 transport error。 ## 工具修复(commit bfe020d) session-stops.mjs 原判定只在独立 type=error 事件时识别 transport-er…

    dust617 · GitHub · Aug 1, 2026
  3. 3

    Which Claude 5? Opus 5 does seem to have diarrhea of the mouth. But Fable 5 hasn't been so bad for me. Or perhaps it is just better at adhering to my guidelines.

    rootusrootus · Hacker News · Aug 20, 2026
  4. 4

    For me on a paid plan, the effort indicator was hidden and the model was 5.5 instant. I had to press the + button to select “think harder” before the dial that allowed Sol medium or high to be selected to show. It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?

    aryehof · Hacker News · Aug 7, 2026
  5. 5

    Scores 75.77% on LiveBench Instruction Following (#5 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  6. 6

    **Bug Description** you flagged something completely innocent and dropped back to opus. it seems pointless to even have fable /model claude-fable-5 **Environment Info** - Platform: darwin - Terminal: ghostty - Version: 2.1.220 - Feedback ID: 5aaf422a-7737-4231-94c4-737c4df0264b **Errors**

    elyochola · GitHub · Aug 2, 2026
  7. 7

    Scores 79.93% on LiveBench Instruction Following (#2 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  8. 8

    Ranks #5 of 144 on LMArena's instruction-following category (Elo 1485), based on blind human preference votes.

    LMArena instruction-following category · Benchmark · Sep 1, 2026
  9. 9

    ### Hermes Web UI Version v0.6043 ### Hermes Agent Version v0.20.3 ### Bug Description 在cli中添加的Google AI Studio模型在web-ui中看不到,如果使用web-ui内置的Google AI Studio添加(base为https://generativelanguage.googleapis.com/v1beta/openai),获取到的模型列表全部带“models/”前缀,如“models/gemini-3.7-flash”,导致对话时gemini返回404 如果手动添加未列出的模型名为“gemini-3.7-flash”则可以正常对话 ### Steps to Reproduce 1、web-ui中“模型”设置页面点击右上角“添加 Provider” 2、“选择 Provider”中选择“Google AI Studio” 3、填入API Key,“默认模型”处点击“获取” 4、点击“添加” 5、新建对话,选择该模型列表中任一模型,输入对话,对话框返回“Error: Gemi…

    hspmanbu · GitHub · Aug 18, 2026
  10. 10

    Model confirmed by the reporter as **`gemini-3.6-flash`**, not `gemini-3.5-flash`. Acceptance updated. Worth noting which way that cuts: `gemini-3.6-flash` is the most expensive of the three configured models and the one #32 is filed against. It handled editor chat code generation with no empty replies — consistent with #32, which only bites at small `max_tokens`, and Continue's chat role does not set one.

    ujjawalmisra · GitHub · Aug 5, 2026
  11. 11

    ## 문제 요약 `gemini-3.6-flash` 무료 티어가 프로젝트+모델 기준 하루 20건 요청으로 제한되어 있는데(`GenerateRequestsPerDayPerProjectPerModel-FreeTier`), 메시지 하나당 LLM을 최소 2번(1차 긴급도 판단 + RAG 재검색 후 2차 최종 판단) 호출하므로 실질적으로 하루 메시지 10개 정도면 한도가 참. 실사용 메신저 봇으로는 감당 불가. ## 에러 로그 ## 해결 방안 (적용함) 채팅 생성 LLM을 Gemini에서 **Groq**(`llama-3.3-70b-versatile`)로 전환. Groq 무료 티어는 분당 요청수 기준이라 하루 단위로 막히는 지금 상황보다 훨씬 넉넉함. - `app/pipelines/shared/llm.py` 신규 — `get_chat_llm()` 공용 팩토리 - `realtime_graph.py`(3곳), `feedback_graph.py`(2곳), `memory_gc_graph.py`(1곳)…

    oh0227 · GitHub · Jul 24, 2026
  12. 12

    ## Problem When proxying Antigravity/Gemini responses to the OpenAI chat-completions format, the `usage` object in a streamed chunk (and potentially non-stream responses) can be emitted **without** `completion_tokens`. Strict OpenAI clients that deserialize `usage` with required fields (e.g. Rust serde in Grok CLI) fail the whole turn: Observed with `gemini-3.6-flash-high` via Antigravity OAuth. Heavy-reasoning responses appear to trigger it: upstream sends a `usageMetadata` chunk containing `p…

    amikai · GitHub · Jul 21, 2026
  13. 13

    **Describe the bug** Returns “No response: RAW” whenever using ‘book chat’ feature. First it does show a streaming response to my prompt, but when it finishes I get the error. **To Reproduce** Using Openrouter and Gemini 3.6 Flash, when I use the ‘Book Chat’. **Expected behavior** It to show me the response to the chat I submitted. **e-reader (please complete the following information):** - Device: Musnap Ocean - OS: Android 14 - KOReader **Additional context** Haven’t had any issues with any o…

    rockettemorton · GitHub · Jul 24, 2026
  14. 14

    Ranks #1 of 144 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.

    LMArena instruction-following category · Benchmark · Sep 1, 2026
  15. 15

    Scores 88.67% on LiveBench Reasoning (#14 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  16. 16

    ### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…

    NStestUser1954 · GitHub · Jul 22, 2026
  17. 17

    ### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…

    NStestUser1954 · GitHub · Jul 22, 2026
  18. 18

    ## Contexto Existem duas implementações independentes do mesmo endpoint de chat ("Consciência Assistida" do AI Lab): - `server.js` (rota `/api/ai-lab-chat`, usada apenas em dev via `npm run dev` — Express local, provedor **OpenRouter**, `OPENROUTER_MODEL` default `openai/gpt-4o-mini`) - `api/ai-lab-chat.js` (Vercel serverless function, provedor **Anthropic** via `@anthropic-ai/sdk`, modelo `claude-opus-4-8`) ## Evidência `server.js`: `api/ai-lab-chat.js`: O system prompt já **divergiu de fato**…

    luksjfernandes-ctrl · GitHub · Jul 26, 2026
  19. 19

    Scores 72.03% on LiveBench Instruction Following (#14 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  20. 20

    Scores 90% on LiveBench Reasoning (#7 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  21. 21

    ### 分支选择 dev-stable ### 模块选择 LLM wiki ### Checklist - [x] 我已经搜索过相关问题,但没有得到预期的帮助。 - [x] 最新版本中该错误尚未修复。 - [x] 请注意,如果您提交的Bug描述缺少相应的环境信息和最小可复现的demo,我们将很难复现和解决该问题,从而降低收到反馈的可能性,甚至该问题将被关闭。 ### 🐞 问题详细描述 ●内置模型来源:后端 Zen 提供的 9 个 free 模型(前端"内置免费模型"列表) 问题汇总 # 前端显示名 后端标识符(已知/待确认) 现象 错误码 错误类型 1 DeepSeek V4 Flash deepseek-v4-flash-free 调用报错 181001 / HTTP 400 BadRequestError: Model is unavailable 2 OX Alpha Free 待确认(推测带-free后缀) 对话无反应 / 无输出 无(无报错) 静默失败(疑似超时或未捕获错误) 3 Muse Spark 1.2 待确认(推测带-free后缀) 调用报错 181001 / H…

    openjiuwen-collaboration-bot[bot] · GitHub · Aug 24, 2026
  22. 22

    Scores 90.51% on LiveBench Reasoning (#6 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  23. 23

    ## Description `ChatBedrockConverse.stream()` crashes with `IndexError: list index out of range` when the model streams redacted reasoning blocks. Amazon Bedrock's `ConverseStream` can deliver reasoning as either `reasoningContent.text` (plain) or `reasoningContent.redactedContent` (encrypted bytes). `_bedrock_to_lc()` handles the first form but has no branch for the second, so it returns an empty list. `_parse_stream_event()` immediately indexes `[0]` into that result and raises. xAI Grok 4.6…

    Great-Root · GitHub · Aug 21, 2026
  24. 24

    添加kimi-k3和grok-4.6后,在使用时会报Could not compact conversation然后终止,没有任何输出。而使用kimi-k2.7-highspeed和glm-5.3则没有这样的问题。请问是不是我配置的有问题呢? ![报错截图](https://files.seeusercontent.com/2026/08/21/kxS5/20260821_102019.png) ![配置截图](https://files.seeusercontent.com/2026/08/21/Wlz2/20260821_102116.png)

    hhr114 · GitHub · Aug 21, 2026
  25. 25

    ## Destination A single lean llama.cpp `llama-server` (HIP, gfx1100) serving Qwen3.8-27B (unsloth `UD-Q3_K_XL`) on the RX 7900 GRE — KV-cache pre-sized to fit in VRAM so it physically cannot OOM, exposed OpenAI-compatible, fronted by the Hollama browser UI as the opencode replacement, and launched under a systemd unit with a hard memory cap. "Done" = a usable model I can chat/code with: no swap, no 27GB RAM spike, correct context actually provisioned. ## Notes - Domain: local LLM serving, llama…

    darylerivers · GitHub · Aug 19, 2026
  26. 26

    ## Why Two workload shapes with opposite requirements are currently served by one always-on GPU: - **Sprinkled interactive traffic** (Discord, chat, vision, classifier, pi turns). Needs a resident model and low TTFT. Served well today by the 4090 running Qwen3.8-27B at 262K context. - **Clumped batch work** (`model-bench` sweeps, bulk classification, embeddings backfills, autonomous queue jobs). Latency-insensitive, preemption-tolerant, and wants a much larger model than the 4090 can hold. The…

    jomcgi · GitHub · Aug 26, 2026
  27. 27

    Scores 72.66% on LiveBench Instruction Following (#13 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.