Open Psychology
Sources
THE AI LAB THESIS / A SHARED EXPLORATION

Teaching AI
to understand
psychology.

A model trained on the field’s knowledge.
Refined through clinical judgement.
Connected to evidence and outcomes over time.

A working paper for Kim & Florian
Based on the model-training concept and meeting notes · September 2026

Start with the summary ↓
00

Executive summary

The whole thesis in two minutes. Each point links to the chapter that explains it.

Build a psychology specialist on top of an existing open-weight model and offer its capabilities to other software through one API. The durable asset would not be a model file. It would be the licensed data, clinician judgement, evaluation and outcome-linked research that accumulate around it.

What we would build

Adapt a strong existing model through continued pretraining on psychology material, fine-tuning on expert tasks, and clinician preference learning, starting with DPO. Candidate base models include Qwen, Mistral, Gemma, MedGemma and Kimi.

Chapters 02–04 →

What would make it different

Training data that captures what clinicians noticed, why they acted and what happened afterwards. That means supervision reasoning, expert preferences and longitudinal outcomes, not just a large pile of transcripts.

Chapters 05–06 →

How it would run

The model is one part of a system. Current guidelines, patient records and scoring rules stay outside the weights. A family of proposed specialist models (reasoning, search, reranking, risk and quality rating) would sit behind one API.

Chapter 07 →

How we would know it works

A clinician-defined evaluation set comes first and is kept out of training. Each training stage is compared against strong general models given the same references and tools, and failure modes are tested adversarially.

Chapter 08 →

Which model to start from

Kimi K2.6 has the highest score on HealthBench-Psych v1 (0.627), but the top five models are statistically tied. K3’s licence likely requires a separate Moonshot agreement for an API business above US$20m annual revenue, so it needs legal review before major investment.

Chapter 09 →

The business

Usage-based API access for clinics, health-record vendors, therapy platforms, researchers and other software companies. Demand, margins and regulatory status depend on the intended use and have not yet been established.

Chapter 10 →

What has to be true

  1. Deeper domain training beats strong general models that are given the same references and tools.
  2. We can identify which data (supervision, preferences, literature, outcomes) drives the improvement.
  3. Expert quality can be defined and measured, including uncertainty and disagreement.
  4. We can secure the data rights, clinical partners and team the programme needs.
  5. The advantage survives real use across settings, languages and populations, at a viable cost.

See the research questions in chapter 11 →

Status: a private discussion document. Nothing described here has been built or validated. The model names describe proposed capabilities, not existing products.

01

The psychology AI lab

The ambition is a specialist intelligence layer that other organisations can build on.

The idea is to take a capable existing AI model and develop it into a psychology specialist. That means teaching it the field’s knowledge, the tasks clinicians perform and the standards experienced practitioners use to judge a response.

Over time, this could become infrastructure for clinics, therapy platforms, health-record software, researchers and other applications that need to understand psychological context. A shared API would let their software call these capabilities.

What could we build if the training data captured what clinicians noticed, why they acted, and what happened afterwards?

What specialisation could add

A better grasp of psychological concepts, more useful case analysis, appropriate uncertainty and behaviour shaped by clinical standards. These are capabilities to develop and measure, not benefits we can assume.

What the company would build

The training corpus, expert feedback, evaluation system, adapted models and the software needed to use them. The long-term asset would be the combination of those parts.

This is the lab thesis from the source document. It does not depend on a claim that general AI providers will stay out of mental health. The research question is whether specialised training and data produce a meaningful, durable advantage.

02

Four depths of specialisation

The original proposal targets a domain specialist, built on an existing model.

LEVEL 1

Reference layer

Prompt + retrieval

Give an existing model instructions and relevant material at the time of a request.

Changes model weights? No

What this buys us

Useful for current guidance, traceable sources and testing the limits of an unchanged model.

The model itself has not learned new behaviour. Quality still depends on the model, the retrieved material and the surrounding controls.

LEVEL 2

Task specialist

Supervised fine-tuning

Train an existing model on expert examples of tasks and good responses.

Changes model weights? Yes

What this buys us

Teaches patterns for case formulation, interviewing, documentation or recognising interventions.

Requires high-quality demonstrations and tests of whether learning generalises beyond those examples.

LEVEL 3

Domain specialist

Continued pretraining + fine-tuning

Continue learning from psychology material, then teach tasks and clinician preferences.

Changes model weights? Yes

What this buys us

The central approach in the memo: build domain familiarity and behaviour into the model itself.

More training and data work. Any benefit over fine-tuning alone must be demonstrated, including whether general capabilities deteriorate.

LEVEL 4

New foundation

Training from scratch

Build a new general language model before developing its psychology capabilities.

Changes model weights? Yes, from the start

What this buys us

Offers the greatest control over the full training process.

An entirely different scale of data, compute and engineering. The proposed lab does not require this starting point.

Weights are the model’s learned numerical parameters. Training changes those parameters. Retrieval supplies information for the current request. Both can be valuable, and neither creates defensibility automatically.

The memo proposes comparing models from Qwen, Mistral and Gemma, including medically adapted MedGemma. Kimi adds another candidate. The right choice depends on the exact checkpoint’s performance, training access, operating cost and commercial terms.

03

The training pathway

Each stage teaches a different thing. Select a stage to explore the full explanation.

STAGE 2 / Domain knowledge and familiarity

Psychology continued pretraining

Continue the model’s broad learning process on a carefully curated psychology corpus. In simple terms, it repeatedly learns to predict the next piece of text from the material before it.

The curriculum could include clinical psychology, psychiatry, counselling, developmental psychology, behavioural science, therapy manuals and case studies. This changes the model’s weights.

Unlike instruction tuning, the material need not consist entirely of questions paired with ideal answers. The goal is richer domain familiarity. Useful clinical behaviour must still be taught and tested.

What goes in
Licensed literature and legally usable domain material, cleaned and reviewed for quality.
What comes out
A psychology-adapted checkpoint, provisionally called PsyBase.

The proposed sequence is base model, continued pretraining, supervised fine-tuning, clinician preferences, then evaluated deployment with references and tools. In practice, evaluation runs throughout. See the full research curriculum. Method references: DPO [5], efficient adaptation [11] and domain-adaptive pretraining [12].

04

What should the model be rewarded for?

Clinical judgement needs a richer definition of “good” than whether a user liked the answer.

Psychological theories of reinforcement describe how people learn through consequences and feedback. Reinforcement learning in machine learning is a mathematical way to update behaviour using a reward signal. They are related ideas, but applying the algorithm does not automatically give a model psychological understanding.

For this lab, the important contribution would be deciding what deserves reward. Experienced clinicians could define the criteria, assess difficult trade-offs and help develop a specialist PsychologyReward model that predicts expert assessments.

ILLUSTRATIVE TRAINING EXERCISE

Two answers can sound helpful. Only one respects the uncertainty.

A fictional case contains limited information about persistent low mood. Which response pattern would you prefer?

Select an answer to see what the training example would capture.

A proposed clinician rating framework

Relationship and autonomy

Empathy, appropriate questioning, therapeutic alliance, respect for choices and avoidance of dependency.

Reasoning and uncertainty

Evidence-based analysis, recognition of missing information and avoidance of over-pathologising.

Method and context

Fidelity to the intended therapy framework, appropriate intervention selection and cultural sensitivity.

Safety and boundaries

Crisis recognition, escalation, professional boundaries and restraint around diagnosis and medication.

DPO: learn directly from comparisons

Train on pairs of preferred and rejected responses. The original paper presents a simpler alternative to the conventional reward-model-plus-reinforcement-learning pipeline. This is the memo’s proposed starting point for preference learning. [5]

RLHF: learn using a reward model

In a conventional reinforcement learning from human feedback pipeline, expert comparisons train a reward model. That model then scores generated responses while the main model learns to improve those scores. It adds a separate optimisation process and opportunities to exploit weaknesses in the scoring system. [5]

For psychology, a high reward score would remain a proxy for quality. It would not establish that a person’s health improved.

RLAIF: use AI feedback, with expert oversight

AI reviewers can compare or critique responses against written principles. Constitutional AI is a precedent for this approach. [6]

The proposed lab could use it to increase coverage while clinicians audit judgements and review serious or ambiguous cases. Agreement between AI reviewers should not be treated as independent clinical validation.

Experts will sometimes disagree. Preserve those disagreements and their context. Do not collapse all therapeutic schools, cultures or safety trade-offs into a single unexplained score.

05

A much richer training corpus

Raw conversations are only one part of the dataset. The strongest material may explain the decisions behind them.

The source document proposes combining domain knowledge with expert practice, measurement and outcomes. Different sources teach different things. A large pile of transcripts is not a substitute for this design.

21 proposed data sources

Data sourceWhat it could teachWhat needs care
Psychology textbooks and curated literatureDomain concepts and vocabularyCopyright, quality and relevance
Research papers and meta-analysesEvidence and its limitsStudy quality, conflicting findings and updates
Treatment guidelinesAccepted practice within a settingJurisdiction and version; retrieve current guidance
Therapy manualsIntervention structure and sequencingRights and fidelity to the intended method
Therapy transcriptsConversation patterns and patient expressionConsent, privacy and contextual interpretation
Supervision discussionsWhy clinicians choose or avoid an actionExpert reasoning, disagreement and data rights
Case formulationsConnecting context, hypotheses and maintaining factorsAvoid teaching conjecture as established fact
Intake assessmentsInformation gatheringMissing context and sampling bias
Mental-status examinationsStructured observationsObserved facts versus interpretation
Progress notesChange across sessionsIncomplete records and authorship differences
Treatment plansHow goals relate to intervention choicesAppropriateness is contextual
Psychometric questionnairesStructured measurements of psychological stateInstrument permissions and correct scoring
Questionnaires with interpretationsHow experts interpret a measurementA score is not a standalone diagnosis
Longitudinal scoresPatterns of change over timeMissingness, context and measurement limits
Session outcome measuresObserved progressAssociation is not evidence of causation
Client feedbackExperience of care and therapeutic allianceSatisfaction differs from clinical benefit
Therapist intervention labelsRecognising specific techniquesReviewer consistency and treatment context
Clinician preference comparisonsWhich response is better and whyRetain rationale and disagreement
Failure and adverse-event examplesRecognising harmful response patternsRare-event coverage and careful handling
Crisis examplesRecognition and escalation behaviourRealistic variation and expert oversight
Expert-reviewed synthetic casesCoverage of rare or difficult situationsGenerated examples can reproduce model errors

Why supervision is particularly interesting

A transcript records what someone said. A supervision discussion can explain why a clinician chose a particular moment, question or intervention.

Partnerships with training institutions could help collect expert-authored cases and deliberation: what was noticed, which hypotheses were considered, what information was missing and what the practitioner would avoid.

ILLUSTRATIVE SUPERVISION NOTE

“I wanted to understand their experience before challenging that belief. We had not yet built enough trust for that intervention.”

This exposes a decision and its rationale. It is a clinician’s account, not proof of the correct action or access to an AI model’s internal reasoning.

Quality hierarchy, synthetic cases and internet data

The proposed emphasis is on expert-created material and curated literature, followed by appropriately licensed clinical records and consented therapy data. Reviewed synthetic cases can fill gaps. Web material should have a clearly defined role.

Forums may help represent how people express distress, but can contain self-diagnosis, harmful advice and substantial selection bias. Treating them as clinical ground truth would undermine the curriculum. Clinicians should define synthetic-case templates and quality standards rather than letting models recursively teach their own mistakes.

Data access is a core research and business problem

Rights must cover the actual use: access, annotation, model training, commercial deployment and any permitted sharing. Removing names alone does not establish anonymity. Plan for data minimisation, security, retention and appropriate governance.

MIMIC-IV-Note is an example of a governed clinical-notes resource, not a ready-made psychology dataset or blanket permission to train commercial models. Its access and use conditions need separate review. [8] Personal health information has special protection under UK data law. [10]

06

Learning about change over time

The larger ambition is to understand a trajectory across weeks and months.

A single exchange shows only a small part of someone’s situation. The dataset imagined in the memo connects their state, a clinician’s assessment and action, the rationale, the response and later observations.

01Context

The person’s starting situation

Symptoms, circumstances, functioning, history and available measurements. Preserve what is known and what is missing.

02Assessment

What the clinician notices

Observations, competing hypotheses and relevant risks. Record the basis for each interpretation.

03Formulation

How the clinician makes sense of it

A structured account of contributing and maintaining factors, with uncertainty and alternative explanations.

04Action

What the clinician chooses

The intervention, its intended goal and the rationale. Include choices to wait, ask more or refer.

05Response

What happens in and after the session

Patient response, experience of the relationship, engagement and relevant changes in context.

06Follow-up

What changes over time

Next-session state, symptom and functioning measures, later outcomes and adverse events. These observations do not by themselves show causality.

This could support research into which intervention sequences are associated with improvement in particular contexts. It could also support evaluation of whether a model tracks change accurately and recognises when its earlier interpretation needs revision.

Outcomes require careful study design. People differ, treatment is not randomly assigned and circumstances change. Directly rewarding symptom-score reduction could teach misleading associations or unwanted shortcuts. Better prediction is also different from knowing which treatment will cause improvement.

Beyond text: potential future inputs

STRUCTURED INFORMATION

Clinical history and measures

Questionnaires such as PHQ-9, GAD-7 or WHO-5, together with medication, sleep, treatment history, functioning and attendance.

EVERYDAY CONTEXT

Brief check-ins

Momentary reports of mood, anxiety, energy, rumination and social connection. Their meaning depends on timing and context.

PHYSIOLOGY

Sleep and activity

Potential signals such as sleep patterns, heart rate and movement. Their usefulness for a particular claim needs evidence.

CONVERSATION SIGNALS

Audio and interaction

Speech pace, pauses and turn-taking. Video or inferred emotional signals would demand especially careful validity, bias and privacy assessment.

These are research directions, not claims that a wearable, voice pattern or brain scan can diagnose a condition. A treatment-simulation “digital twin”, mentioned in the meeting notes, is a longer-term hypothesis requiring its own evidence.

07

The model is one part of the system

Learn durable concepts in the model. Supply changing information and precise tools when it runs.

What belongs in learned behaviour

Psychological language and concepts, therapy frameworks, task skills, conversation patterns and habits such as acknowledging uncertainty.

What belongs outside the weights

Current guidelines, local policy, medication references, authorised patient records, scoring rules and evidence citations.

A clinical guideline changes

Retrieve the current source

Keep changing guidelines in an approved reference system with dates and citations. The model can use them at the time of the request without waiting for retraining.

A retrieval update changes the reference available to the model. It does not guarantee the model interprets it correctly.

Training and a controlled workflow can reinforce each other

The meeting notes contrast fine-tuning with a deterministic “harness”: software that constrains the sources, tools and steps a model can use. A knowledge graph could make some retrieval paths explicit. Neither fixed rules nor a graph make all generated language deterministic, complete or clinically correct.

The lab could combine both approaches. Training develops capability and behaviour. The surrounding software enforces access, routes tasks, checks defined conditions and records evidence. Each contributes something different.

A family of specialist models

The original vision includes several components behind one service. These names describe proposed capabilities, not models that already exist.

PsychologyLMReasoning and conversation

The adapted language model handles psychological concepts, case analysis and dialogue within its validated scope.

PsychologyEmbedFinding related concepts

Turns text into representations that help search for relevant cases or literature. Similarity needs domain-specific testing.

PsychologyRerankerChoosing useful evidence

Reorders retrieved material so the most relevant sources reach the model.

PsychologyRiskConstrained risk detection

Flags defined risk signals for an escalation process. False negatives, false positives and operational follow-through all matter.

PsychologyRewardPredicting expert ratings

Estimates quality against clinician-defined criteria. It remains a fallible proxy, subject to expert audit.

Partner applicationShared APIModels + references + toolsChecked response with sources
08

A psychology-specific research curriculum

Validation should shape the programme from the beginning, not appear after the model is trained.

  1. 0

    Create the evaluation set

    Clinicians define tasks, rubrics and failure cases. The memo suggests an eventual 2,000–10,000 cases as a planning range, not a validated sample-size requirement. Keep evaluation cases out of training.

  2. 1

    Compare base models

    Use the same cases and documented configurations. Compare open-weight candidates and strong hosted models, including simple prompt-and-retrieval baselines.

  3. 2

    Continue psychology pretraining

    Build a licensed, curated corpus and train a domain checkpoint. Measure what improves and what degrades compared with the base.

  4. 3

    Teach expert tasks

    Fine-tune on demonstrations of the work. Check generalisation to cases, institutions and populations absent from training.

  5. 4

    Learn clinician preferences

    Collect response comparisons and rationale. Test DPO and assess whether the model improves on clinically meaningful criteria.

  6. 5

    Test failure modes

    Conduct adversarial evaluation for crises, delusions, harmful advice, dependency and professional boundaries. Red-team the surrounding software too.

  7. 6

    Study use with practitioners

    Use appropriately governed studies or deployments with structured feedback. This is an evidence stage, not a commitment to one initial product.

  8. 7

    Study longitudinal outcomes

    Where permissions and study design allow, connect decisions to later observations. Use this to refine research and evaluation without confusing association with treatment effect.

What the evaluation needs to cover

Capabilities and therapeutic approaches

Assessment, case formulation, documentation, longitudinal reasoning and measurement. Include cognitive behavioural therapy (CBT), acceptance and commitment therapy (ACT), motivational interviewing (MI), psychodynamic approaches and other relevant methods.

Risk, context and equity

Suicidality, mania, psychosis, abuse, eating disorders, medication, minors and relationship boundaries. Test relevant languages and cultural contexts, missing records and contradictory information.

Human experts should assess clinically important errors even when automated judges help with scale. Compare stages independently so we know whether improvements come from domain pretraining, examples, preferences or the surrounding tools.

A benchmark measures defined tasks under evaluation conditions. A practitioner study measures performance in its setting. Claims about safety or patient benefit require evidence appropriate to that intended use. Neither a model name nor a high average score establishes universal clinical validity.

09

Kimi: evidence and commercial terms

A useful model-selection case study within the larger lab strategy.

What the benchmark tells us

HealthBench-Psych v1 compares 20 models on 610 selected mental health conversations. Clinicians validated case selection, and three AI judges graded answers against physician-written criteria. Kimi K2.6 has the highest mean score. The top five models are statistically tied. K3 is a different release and scores lower here. [1]

Kimi K2.60.627Highest point estimate in v1
Kimi K30.568Separate model and licence

These scores reflect rubric performance, not treatment effectiveness. The live repository now lists 611 cases and 23 models. The figures here match the supplied post and paper v1. [2]

Explore all 20 results and confidence intervals
ModelMean score95% confidence interval
kimi-k2.60.6270.606–0.647
gpt-5.50.6240.605–0.643
claude-opus-50.6200.598–0.641
grok-4.50.6120.589–0.634
gpt-5.6-sol0.6100.589–0.629
claude-fable-50.5910.569–0.613
gemini-3.6-flash0.5780.555–0.602
kimi-k30.5680.545–0.591
deepseek-v4-pro0.5540.531–0.577
qwen3.7-plus0.5520.529–0.577
mistral-large0.5440.517–0.569
deepseek-v4-flash0.5380.516–0.561
claude-sonnet-50.5330.510–0.556
gemini-2.5-pro0.5270.504–0.553
gpt-4.10.5120.486–0.537
gemini-2.5-flash0.4570.431–0.484
qwen3-8b0.4460.419–0.471
claude-haiku-4.50.4410.416–0.463
mistral-small0.3630.336–0.391
gpt-3.5-turbo0.1760.149–0.201

Higher is better on this benchmark. Scores are on a 0–1 scale. The first five models form the statistically tied leading group. Source: paper v1, Table 1. [1]

K3’s custom licence, in plain English

Open weights are downloadable parameters. Their licence still governs commercial use. K3 permits deployment, modification, fine-tuning and distribution subject to conditions, including notices and legal compliance. [3]

Model API business
If the licensee or an affiliate operates Model as a Service and combined revenue exceeds US$20m over any consecutive 12 months, a separate Moonshot agreement is required before commercial use, including derivatives.
Meaning of the service
Third parties have meaningful control over inputs, parameters or training data. Feature-embedded end-user products and simple relaying to others’ hosted models are excluded.
Branding threshold
Products exceeding 100m monthly active users or US$20m monthly revenue must prominently display “Kimi K3”.
Exceptions
Sections 2 and 3 exempt internal use without third-party access to outputs or capabilities, and use through Moonshot’s official products or certified inference partners.

Our interpretation: the proposed specialist API likely fits Model as a Service. Future agreement terms are unspecified. That dependency matters before significant K3-specific investment. [3]

K2.6 has separate terms. Its Modified MIT licence permits commercial use with notices and a branding requirement above 100m monthly active users or US$20m monthly product revenue. The reviewed text does not contain K3’s annual-revenue agreement clause. [4]

Commercial use appears feasible in principle. Model quality, hosting cost, data rights and the exact business configuration still determine suitability. This is a preliminary licence reading, not a legal opinion. Obtain counsel’s review of the pinned release and intended deployment.

10

The company behind the models

The business thesis is to sell specialist capabilities through an API, supported by assets that improve with research.

A customer’s software would send a request to the service and receive a response. Possible capabilities include understanding a case, formulation, measurement, evidence retrieval, summarisation, supervision support and defined risk detection. More consequential functions, such as assessment or intervention recommendations, require stronger evidence and a carefully defined regulatory scope.

CONCEPTUAL DEVELOPER INTERFACE
/understand   /formulate   /assess   /reflect
/intervention   /risk   /measure   /retrieve
/summarise   /supervise

One connection to a family of psychology capabilities. These are proposed functions, not an existing service.

Potential customers

Clinics, electronic health-record companies, therapy platforms, mental health applications, research groups, employee-assistance providers and other software businesses.

These are possible buyer groups. Demand, integration needs and willingness to pay remain open commercial questions.

Possible pricing

Usage-based charges, including per-token billing, alongside enterprise support or deployment arrangements. A token is a small piece of text processed or generated by a model.

Margins must account for inference, hosting, clinical evaluation, data acquisition, training, security and support. Downloadable weights do not mean free operation.

What could become difficult to reproduce

  • Psychology corpus: carefully acquired and curated material, with clear rights.
  • Expert judgement: supervision rationale, demonstrations and preference data.
  • Evaluation: independent cases, meaningful rubrics and documented failure analysis.
  • Outcome-linked research: longitudinal records with suitable permissions and study design.
  • Model and system work: adapted checkpoints, adapters, specialist models and reliable deployment.

The potential advantage is cumulative. Owning a fine-tuned file alone does not establish a durable business. Rights in the base model, the data and the resulting work must also be clear. The source memo’s “you own” language should therefore be read as potential company assets subject to the relevant agreements.

Clinical responsibility and regulation

Generating a summary and recommending treatment can carry different intended purposes and risks. UK MHRA guidance considers both purpose and functionality when determining medical-device status. Supplying an API or keeping a clinician in the loop does not automatically remove obligations. [9]

Define the roles of the model supplier and downstream application provider. Evidence, monitoring, privacy and responsibility for escalation need to fit the actual deployment and jurisdiction. No specific UK, EU or US classification has been established for this concept.

11

The research questions that define the lab

These are questions about the overall model programme, rather than a prescribed first product.

01

Does deeper domain training add value?

Compare continued pretraining plus fine-tuning against fine-tuning alone and strong general models with the same references and tools.

02

Which data creates the largest improvement?

Test the contribution of supervision, preferences, literature and longitudinal information. The memo’s emphasis on clinician reasoning is a hypothesis worth measuring.

03

Can we define and measure expert quality?

Build rubrics that preserve uncertainty and disagreement while detecting serious failures. Establish where automated judging helps and where it needs human review.

04

Can we acquire the rights and partnerships?

Research institutions, training organisations and clinical partners could contribute expertise and governed data. A credible clinical and machine-learning team is central to the programme.

05

Does the advantage survive real use?

Investigate transfer across settings, languages and populations, plus deployment cost and customer demand. A strong internal benchmark is one part of that answer.

The long-term vision is a clinical learning system:
one that connects expertise, evidence and change over time.

A plain-English glossary

Open a term whenever the technical language gets in the way.

Weights

The numerical parameters learned during training. They influence how the model processes inputs and generates outputs.

Open weights

A release that lets people access model parameters. Use, modification and commercial rights depend on its licence.

Continued pretraining

Further training of an existing model on a chosen body of material, often using next-token prediction.

SFT

Supervised fine-tuning: learning from examples that pair an instruction with a desired response.

DPO

Direct Preference Optimisation: training with preferred and rejected answers to shape behaviour.

RLHF / RLAIF

Reinforcement learning using feedback from people / AI. The source of feedback and the training method are separate design choices.

Adapter / LoRA

A smaller trainable addition to a model. LoRA is one method for adapting a model without updating all its original weights.

RAG / retrieval

Looking up relevant material and providing it to the model when answering, rather than relying only on learned parameters.

Harness

The surrounding software that controls workflow, permissions, tools and checks.

Embedding / reranker

An embedding helps represent meaning for search. A reranker reorders retrieved results by relevance.

Held-out evaluation

Testing on cases excluded from training. It estimates performance beyond the examples used to teach the model.

API / inference

An API lets software request a capability. Inference is the process of running a trained model to produce an output.

Sources and scope

  1. HealthBench-Psych, paper v1

    Flathers, Torous and colleagues, August 2026. Benchmark methodology, scores and limitations.

  2. HealthBench-Psych code and data

    The live repository describes a newer release. The comparison here preserves the supplied screenshot and paper v1.

  3. Kimi K3 License

    Moonshot’s official terms, including the Model as a Service condition.

  4. Kimi K2.6 licence

    Official Modified MIT licence for K2.6. Separate from K3.

  5. Direct Preference Optimization

    Rafailov and colleagues. The method behind learning directly from preferred and rejected responses.

  6. Constitutional AI

    Anthropic’s research on training with principles and AI feedback.

  7. MedGemma 1.5 model card

    Google’s description of medical adaptation, intended use and limitations.

  8. MIMIC-IV-Note

    An example of a governed clinical-notes resource. Its existence does not establish rights for this project.

  9. MHRA digital mental health guidance

    UK guidance on intended purpose, functionality and medical-device qualification.

  10. ICO special category data guidance

    UK guidance covering personal health information and related protections.

  11. LoRA: Low-Rank Adaptation of Large Language Models

    Hu and colleagues. Efficient adaptation with a smaller set of trainable parameters.

  12. Don’t Stop Pretraining

    Gururangan and colleagues. Research on adapting language models to domains and tasks, not evidence for this proposed clinical system.

The main structure comes from the supplied ChatGPT concept document on psychology model training. The Kim × Flo meeting notes inform the lab ambition, the discussion of training versus runtime controls and the need for credible research partners.

The architecture, dataset priorities, model names such as PsychologyReward, potential customers and commercial advantage are proposals. They are not results from an implemented system. Method references establish what techniques do, not that this particular lab approach will succeed.

Benchmark figures preserve the supplied paper v1. Kimi terms were reviewed from official licence files on 28 September 2026. Other linked sources support method explanations and regulatory context. The page is self-contained and works offline; opening external sources requires an internet connection.