MyPain. Clinician-supervised pain assessment, on the web and in WhatsApp
Someone living with chronic pain can describe what they are feeling in a conversation rather than a long form. MyPain cannot invent a condition to explain it, because the set it draws from is a table clinicians control. A clinician can open any assessment afterwards and follow how it got there.
424 curated conditions 17 attributes each 6,725 indexed sections Clinical decision support
Pain assessment a clinician can stand behind.
The founding clinician had spent years watching patients arrive at appointments carrying answers they had got from general-purpose assistants, some of them useful and a lot of them not. MyPain was briefed as a conversation that helps a person with chronic pain put words to what they are experiencing, only ever suggests a condition the clinical team has curated, checks red flags against what the patient actually described, and can be read back afterwards by the clinician who has to stand behind it.
This is decision support under clinical supervision rather than a regulated medical device, and it was designed that way from the first architecture session. Before anyone wrote product code, OpenKit produced a written specification and a go or no-go decision the client signed. What followed was the first working version of the service: the conversation itself, the search behind it that answers a patient’s questions, the goal-setting side, and both surfaces the service runs on, the web and WhatsApp.
- 01
One assistant the patient talks to, with diagnosis, analysis and knowledge search working behind it, following the order a GP consultation takes.
- 02
A design where the only conditions the service can return are the ones in the clinical set, so the safety boundary sits in the data.
- 03
A simulator that runs the whole condition set through the product in two different ways of speaking, and scores the results.
What a patient meets.
The same conversation, carried into WhatsApp.
The assessment happens on the web, and the goal that comes out of it lands on the surface the patient already has open. Both sides read one record, so picking the thread up later does not mean starting again.
Illustration of the hand-off, not a screenshot.
MyPain
That’s your assessment complete. I’ve put what we agreed into a goal card you can keep.
SMART goal
- Specific
- Set from what you described, in your words
- Measurable
- Tracked against the same scale each time
- Review
- At the interval you chose
If anything changes, message here. It picks up where we left off.
Got it, thanks
The safety requirements the product had to meet.
- 01
Conditions come from a clinician-curated set
Ask a general assistant about a symptom and a rare, vivid condition can appear in the answer with nothing behind it. In a pain product, an invented condition is a safety problem, and no amount of instruction in a prompt makes it structurally impossible.
- 02
Emergencies are checked against what the patient described
A common condition and a rare emergency can surface together for the same description, presented with the same weight and no sense of which one to act on. A patient reading that has been given more anxiety and less direction than they started with.
- 03
Intake runs as a conversation
Patients drop out of long forms, and the ones living with chronic pain are the least likely to finish one. The conversation had to meet them in their own words, on the surface they already use, rather than asking them to translate themselves into fields.
- 04
Every conversation can be reviewed
A clinician cannot stand behind a tool they cannot inspect. Every conversation has to be reviewable end to end, including which candidate conditions were ruled out and at which question they went.
How the assistant is bounded.
Every message the patient sends goes to one agent, and that agent decides whether to call the diagnosis agent, which filters the clinical table with a query; the analysis agent, which ranks and presents what came back; or the knowledge search, which retrieves educational content. It then puts the answer into language a patient can use. What it cannot do is originate anything, because the boundary is enforced in the data itself rather than in a paragraph of prompt asking the model to behave.
- 01
The patient-facing agent cannot generate a diagnosis, and cannot expand beyond what the diagnosis agent hands it.
- 02
The diagnosis agent queries a curated table. A condition that is not in that table cannot be returned by any route.
- 03
Emergency indicators are checked against what the patient actually described before the conversation is allowed to raise one.
How candidate conditions are narrowed.
Matching a patient’s words against the conditions that seem to fit drops valid diagnoses. A patient saying "shooting pain" where the clinical set says "radiating pain" gets filtered out of a diagnosis that was valid for them, and neither they nor the clinician ever finds out it was considered. So the system works the other way round: it progressively removes what definitely does not apply and keeps everything else in play.
Unknowns are preserved rather than treated as mismatches, catch-all values are never excluded, and age and duration act as soft signals that change the ranking instead of hard filters that delete candidates. As the conversation narrows, the surviving set narrows with it, and the record of what went at which question is kept.
- 01
"Pain in my head" leaves 45 candidate conditions in play.
- 02
"Behind my eyes" narrows the same conversation to 12.
- 03
Every condition that left the set did so for a recorded reason.
How it was tested.
conditions in the curated set, every one driven through the simulator rather than a sample of the easy ones.
modes of patient speech tested: direct factual answers, and the hedged natural language people actually use.
turns, past which a conversation is flagged as taking too long to arrive anywhere useful.
indexed sections behind the educational answers, searched semantically and by keyword together.
The scenario simulator plays a patient for each condition twice over. In the first mode it answers accurately and directly, the way a test fixture would. In the second it hedges the way people do, saying "about a week" and "I think it started after". A separate model scores each run on whether the right conditions surfaced, whether emergencies were caught, how efficiently the questions got there, and whether the output stayed inside its rules.
What the simulator changed.
For one symptom presentation the conversation raised both a very common condition and a rare emergency at the same time. A reviewing clinician would refuse to sign that off.
The fix went into the sequence of the conversation itself. Clinical history moved to the front, new history questions went in to establish whether the pain is new or already long understood, pain already lasting beyond three months now rules out the non-conditional emergencies, and the remaining emergency indicators became conditional on criteria checked against what the patient described. The conversation now takes a history before it raises anything urgent.
What the clinical team controls.
- 01
The clinical set is edited by clinicians
Conditions, their attributes, their management plans and their emergency criteria are curated through the admin dashboard by the clinical team, who alone maintain the set. When the clinical position changes, the system changes with it.
- 02
Every journey can be opened and read
Conversation logs, the filter progression that narrowed the candidates, the model interactions and the timestamps are all captured, so a single patient journey can be read back in full.
- 03
Anomalies are graded by severity
A wrong suggestion or a missed emergency is graded critical, a false emergency high, and an over-long conversation or a skipped clinical question medium. The report tells the team what to look at first.
- 04
Scope was bounded on clinical risk
Accuracy was prioritised for the common presentations, with uncommon and rare conditions surfaced alongside guidance that a medical opinion is recommended. That boundary was written into the scope rather than discovered later.
The build
- React on the front end
- Supabase PostgreSQL and Edge Functions
- Twilio for the WhatsApp surface
- Langfuse prompt and trace platform
- Model-agnostic agent layer (Mistral / Gemini launch models)
OpenKit certifications
- ISO 27001
- ISO 9001, UKAS-accredited
- Cyber Essentials
Controls on this project
- Operates to UK GDPR
- UK data residency
- Clinical decision support, not a regulated medical device
More of the work.
Find your first workflow.
We start with a conversation, audit where AI actually pays back, and build the first automation into how your team already works. We reply within one working day.
Get in touch Governance & compliance