What the benchmark tells us
HealthBench-Psych v1 compares 20 models on 610 selected mental health conversations. Clinicians validated case selection, and three AI judges graded answers against physician-written criteria. Kimi K2.6 has the highest mean score. The top five models are statistically tied. K3 is a different release and scores lower here. [1]
Kimi K2.60.627Highest point estimate in v1
Kimi K30.568Separate model and licence
These scores reflect rubric performance, not treatment effectiveness. The live repository now lists 611 cases and 23 models. The figures here match the supplied post and paper v1. [2]
Explore all 20 results and confidence intervals
| Model | Mean score | 95% confidence interval |
|---|
| kimi-k2.6 | 0.627 | 0.606–0.647 |
|---|
| gpt-5.5 | 0.624 | 0.605–0.643 |
|---|
| claude-opus-5 | 0.620 | 0.598–0.641 |
|---|
| grok-4.5 | 0.612 | 0.589–0.634 |
|---|
| gpt-5.6-sol | 0.610 | 0.589–0.629 |
|---|
| claude-fable-5 | 0.591 | 0.569–0.613 |
|---|
| gemini-3.6-flash | 0.578 | 0.555–0.602 |
|---|
| kimi-k3 | 0.568 | 0.545–0.591 |
|---|
| deepseek-v4-pro | 0.554 | 0.531–0.577 |
|---|
| qwen3.7-plus | 0.552 | 0.529–0.577 |
|---|
| mistral-large | 0.544 | 0.517–0.569 |
|---|
| deepseek-v4-flash | 0.538 | 0.516–0.561 |
|---|
| claude-sonnet-5 | 0.533 | 0.510–0.556 |
|---|
| gemini-2.5-pro | 0.527 | 0.504–0.553 |
|---|
| gpt-4.1 | 0.512 | 0.486–0.537 |
|---|
| gemini-2.5-flash | 0.457 | 0.431–0.484 |
|---|
| qwen3-8b | 0.446 | 0.419–0.471 |
|---|
| claude-haiku-4.5 | 0.441 | 0.416–0.463 |
|---|
| mistral-small | 0.363 | 0.336–0.391 |
|---|
| gpt-3.5-turbo | 0.176 | 0.149–0.201 |
|---|
Higher is better on this benchmark. Scores are on a 0–1 scale. The first five models form the statistically tied leading group. Source: paper v1, Table 1. [1]
K3’s custom licence, in plain English
Open weights are downloadable parameters. Their licence still governs commercial use. K3 permits deployment, modification, fine-tuning and distribution subject to conditions, including notices and legal compliance. [3]
- Model API business
- If the licensee or an affiliate operates Model as a Service and combined revenue exceeds US$20m over any consecutive 12 months, a separate Moonshot agreement is required before commercial use, including derivatives.
- Meaning of the service
- Third parties have meaningful control over inputs, parameters or training data. Feature-embedded end-user products and simple relaying to others’ hosted models are excluded.
- Branding threshold
- Products exceeding 100m monthly active users or US$20m monthly revenue must prominently display “Kimi K3”.
- Exceptions
- Sections 2 and 3 exempt internal use without third-party access to outputs or capabilities, and use through Moonshot’s official products or certified inference partners.
Our interpretation: the proposed specialist API likely fits Model as a Service. Future agreement terms are unspecified. That dependency matters before significant K3-specific investment. [3]
K2.6 has separate terms. Its Modified MIT licence permits commercial use with notices and a branding requirement above 100m monthly active users or US$20m monthly product revenue. The reviewed text does not contain K3’s annual-revenue agreement clause. [4]
Commercial use appears feasible in principle. Model quality, hosting cost, data rights and the exact business configuration still determine suitability. This is a preliminary licence reading, not a legal opinion. Obtain counsel’s review of the pinned release and intended deployment.