Perspectives

AI improved everything except the outcome

White text on a bordeaux field reading everything improved except the outcome, above the Nexus Recognita logo.

István Borbíró · [publish date] · 3 min read

Sixteen primary care clinics in Kenya ran one of the most careful AI trials published this year (Agweyu et al. 2026). GPT-4o read along with the consultations of 9,691 patients, flagging concerns in real time. On the expert panel's review of two thousand encounters, the documented care improved in every domain scored: appropriate diagnoses, comprehensive notes, adequate treatment plans.

Then the rest of the trial. The primary outcome, treatment failure within fourteen days, was 2.2% with the tool and 2.0% without it. The concrete practice measures told the same story: no difference in correct antibiotic use, in antimalarial prescribing, in the diagnosis and treatment of hypertension, in the detection and referral of child malnutrition. Patient satisfaction identical. No time saved either; the median consultation was eleven minutes in both arms. The direction of effect favoured the tool, and the authors write that any true benefit is probably modest, smaller than the trial could detect.

The notes got better. The patients did not.

Three weeks earlier, a different field published its own version (Strömberg et al. 2026). Researchers followed 26,811 secondary school students for thirty months as generative AI arrived in their homework. Homework scores rose 18% and completion time fell 30%. Within six months, closed-book exam scores fell 20%. Entrance exam scores fell 18% and 24%, with the full penalty emerging over two years.

Read carefully, even the homework gain empties out. The paper treats unusually fast completion with high homework scores as the signature of outsourcing, and the learning losses concentrated in the roughly 80% of AI users who fit it. Students who kept their working time steady while using the tool lost almost nothing. The saved time was the removed effort, measured.

The students saved time. The learning went with it.

Better notes and faster homework are not better outcomes. The reason is worth more than the results.

AI improves what is there. It was given the note, so the note got better. It was given the homework, so the homework got faster. What it was never given is the purpose those artefacts serve. A consultation note exists so that a decision gets made well. Homework exists so that effort builds something in a student. The tool polished what was in front of it and left the purpose where it found it. In the classroom, the purpose was the effort itself, so the polish removed it. In the clinic, the purpose sits in the decision and the plan that follows, and the trial suggests the record was not what held them back.

AI improves what is there. Only purpose turns better details into better outcomes.

Neither result is an argument against the tool. It is an argument about what a tool can be given. Purpose is held by people: the clinician who reads a flag against the patient in front of them, the team that decides whether a recommendation changes the plan, the teacher who protects the effort that makes homework worth doing. Neither trial wired the tool to that layer. The Kenya authors say as much in their own register, asking openly what the right primary outcome is for a general-purpose technology. It is the honest frontier, and it is not a modelling question.

So one question is worth carrying into the next study headline, vendor slide, or pilot review: what purpose was this tool given, and did anything change there? In clinical terms: did the decision change, or only its record? In cancer care, the layer that holds purpose between people has a name: it is the coordination the team runs on, and it remains largely uninstrumented. Which is why it keeps disappearing between a trial's process measures and its endpoint.

Both papers reward reading in full. They are the clearest demonstration yet that what AI improves and what care is for are measured in different places.

Reference: Agweyu A, Mwaniki P, Menon V, et al. Nature Medicine. 2026. doi:10.1038/s41591-026-04503-6.
Reference: Strömberg D, Lei V, Wu Y. CEPR Discussion Paper 21577. CEPR Press. 2026.