Ask what humans still bring once AI does the work, and the same words come back from every corner of the discourse: judgment, taste, agency. Paul Graham says that when anyone can make anything, the big differentiator is what you choose to make, and the president of OpenAI compressed it further: taste is a new core skill. The consensus is real, it formed fast, and it is pointed the right way. It is also empty. Every word in it is a mystery word, and the words survive precisely because they are mysterious. Nobody can test for what they name. The tests that carry the name measure whether you know the approved answer to a scenario that has one; the judgment everyone suddenly prizes is what you need when there is no key, and no instrument reaches it. Nobody can teach taste. Nobody can tell you whether they have either one or merely believe they do, and the terms flatter whoever claims them, which even their boosters have started admitting out loud. We have named the location of the human edge and left the edge itself unnamed, which means we cannot find the people who have it, cannot build it in the people who don’t, and cannot notice when we’re rewarding the word instead of the capacity.
I think I know what is under the words, because I watched it work nearly three decades before it had a market. It is older and narrower than judgment. It is the same skill whether the problem is a tractor or a hallucinating model. And every instrument we use to find talent, including the consensus vocabulary itself, is built to miss it.
In the late 1990s, before any of the titles, I answered phones for Compaq to pay my way through college. Technical support, the bottom rung — the job that funds a political science degree and asks nothing about your résumé. One afternoon I was pulled off the floor because a call of mine had been audited, someone had listened in, and my manager wanted to understand how it had gone the way it did.
The customer had called because her son’s tractor was fuzzy. That was the whole report. Yesterday the tractor was clear; today it was fuzzy. She could tell me nothing else, because as far as she was concerned there was nothing else to tell.
There was never any question of a literal tractor — she had dialed a Compaq help line, so whatever was wrong was wrong with her Compaq. The problem was that she didn’t have the words for that. The machine was a foreign country to her, and people describe foreign countries in the vocabulary of home. Her home was a farm family: a son she was proud of, a grown man with a farm of his own, whose tractor she was proud of too — proud enough that a photograph of it was the desktop wallpaper on her computer. Yesterday the photograph was clear; today it was fuzzy. She reported the change honestly, in the only frame she had. So the ten seconds were spent translating, not discovering. Strip her report of everything that belonged to her frame alone — the son, the pride, the tractor as a machine in the world — and what remains has to be true in any frame: something the machine displays was crisp and is now degraded. That remainder has a name on the machine’s side of the border — a resolution fault — and a short list of causes in a known order. The setting first; no change. Then the driver. I walked her into safe mode, had her strip the display driver out, restart, reinstall, restart again. The tractor was clear. She thanked me for fixing the tractor, which, in every sense that mattered to her, I had.
I knew those machines cold, the drivers, the safe-mode routine, all of it, and none of that knowledge is what cracked the call. It could not, because until the problem was rebuilt there was nothing for it to act on. What cracked it was seeing the shape under the words, image and not object, blurry and not fuzzy, resolution and not tractor, and matching that shape to a structure I already carried. The content was the easy part, and I had all of it. The structure was the hard part, and it was the whole of the job.
If you want to call what happened on that call judgment, I won’t argue. But notice that the word explains nothing. It renames the mystery. What happened on that call was a mechanism, it has parts, each part can be watched, and each part can be trained.
The four gates
Strip away the surface of that call and there is a mechanism underneath. It runs through four gates, and each has to be cleared before the next one matters. I did not know I was passing through them then. What convinced me the gates are real is twenty years of watching people fail them, because the failures do not look alike. A person missing one gate fails differently from a person missing another, and telling them apart is what the mystery words cannot do. Judgment names the whole; the gates name the part that broke, which is the difference between a compliment and a diagnosis.
The first is a library. Before you can read a problem’s structure you need a stock of structures to read it against, and that comes only from real depth in something. This is the part the taste discourse gets closest to and still misses: taste is treated as a sensibility you acquire by exposure, when the working version is a library you build by depth. It is not the dabbler with a little of everything who makes the leap; it is the person with genuine expertise in at least one place, because expertise is where you learn what holds a structure up and what merely hangs on it, which variables tend to decide outcomes and which are decoration, and the nerve to throw out what you are being told and rebuild it. Matthew Crawford spent a book, Shop Class as Soulcraft, making this case for the mechanic’s kind of knowing: the seeing that only long acquaintance with particulars can buy. The library is that claim, carried out of the shop. The customer told me about a tractor, and depth is what made her report worth mining instead of dismissing: I knew the world she was speaking from, and I carried the machine-side shapes her signal could land in.
The second is the cut. You take the problem as given, over-described and noisy, and you reduce it to the few variables that actually determine the outcome. Everything else, however loud, goes in the pile marked does not matter. This is the hard step. It is the discipline of separating the important from the merely present, the immutable from the optional, structure from content, the variable that solves the equation from the twenty that dress it up. The cut on that call discarded the furniture of her frame — the son, the pride, the tractor as a machine in the world — and kept every bit of signal her words carried. What survived belonged to no frame, which is exactly what made it usable in mine.
The third is recognition. You match what the cut left standing to a shape in your library, and you build toward it. The transfer research calls this step analogy, and analogy is the wrong word. It implies a reach across a gap, a source in some other field. My own language for it, years before I read any of the research, was closer to the truth: I wanted people to hold the why, not the what or the how, because only the why travels. When I built “resolution fault” out of “fuzzy tractor,” I was not reaching anywhere. I was recognizing a shape I had learned in my own trade, which is what depth is for. It was one of the small number of structures I had seen before, surfacing under a new set of clothes. But recognizing it did not consult the trade. Matching asks whether a shape fits, not where you got it. Whether the clothes come from my own trade or from one a thousand miles away changes nothing about the act. Recognition is distance-blind. That fact will matter later, because the accounts that came closest to naming this capacity attributed it to distance, the property that makes the act visible rather than anything that makes it work.
The fourth is the check on the other three. You try to break your own build before you rely on it. The reason is structural, not psychological. A build is produced by matching, and matching tests for fit against the library, not truth against the world; a wrong match can be perfectly coherent, because coherence is the only thing construction ever checked. Worse, whatever read produced the error will approve it on inspection, since the blind spot that made the mistake is the same blind spot examining it. So the only defense is to change roles, to stop being the builder and attack the thing as if someone else had made it, to go looking for the reason it is wrong before the world finds it for you. Crawford retells a story of Robert Pirsig’s, from Zen and the Art of Motorcycle Maintenance, that shows what happens otherwise. Pirsig’s motorcycle kept seizing, and the shop matched the symptom to the standard answer, tappets, and adjusted them; when the seizures continued they kept working the wrong build, chiseling covers, shearing bolts, wrecking the machine a little more each visit. The real cause was a sheared twenty-five-cent pin starving the engine of oil, which Pirsig eventually found himself. The mechanics were not unskilled. They had a match, the match felt finished, and nobody tried to break it, so the world broke it for them at the engine’s expense. Skipping the break is easy, because every build arrives feeling finished. And the machine has just raised the price of skipping it: an AI produces exactly this, a coherent build delivered in a confident voice, at any volume you ask. The people worth the most in that partnership are the ones who reach for the break first, who treat the fluent answer as a candidate rather than a verdict.
Library, cut, recognition, break. Depth to hold the structures, the discipline to strip a problem to the ones that matter, the recognition that a shape has appeared before, and the doubt to test the match before trusting it. That is the mechanism. It is not a sensibility, not a vibe, not a talent for connecting distant fields. It is a way of operating one layer below the field, and every one of its four parts can be observed, which is what makes it different in kind from the mystery words. You cannot test for taste. You can test whether a person can strip an over-described problem to its determining variables, because I have watched it done and watched it fail, one gate at a time.
I should admit that I misnamed this mechanism myself, from inside it, for years. The word I used was translation — I spent seasons of my career sitting between lawyers and technologists, taking a report built entirely in one profession’s vocabulary and rebuilding it in the other’s, and translation was the nearest term for the outside view of what I was doing. It was wrong in an instructive way. Translation implies a dictionary, a mapping between surfaces, and there is no dictionary entry from tractor to display driver, none from a privacy statute to a deletion pipeline. What ran under my word, every time, was the mechanism: reduce the report to what must be true in any frame, find the shape both frames share, rebuild where the fix lives. I was performing the skill daily and still named it by its costume — which is the whole predicament of the current consensus, compressed into one biography. The mystery words are not lies. They are what everyone reaches for when they watch the mechanism work and grab the nearest term for the outside view.
The obvious objection is that this is intelligence with extra steps, or expertise with a new name. Gary Klein spent decades documenting how experts under pressure recognize a situation and act on it without weighing options, which is gate three seen from the inside, and ability testing has claimed this territory for a century. I am not claiming the gates float free of either one; someone who tests well will clear them more often. The difference is what the score leaves out. An ability test reports that a person arrived at a good answer. It cannot tell you whether the one who missed had no library to match against, failed to strip the problem, matched it to the wrong shape, or built soundly and never checked. Those are four different failures with four different repairs, and unlike a general ability score, every one of them can be trained.
Why the shapes repeat
The mechanism only works because of a fact about problems, and it is worth saying plainly, because it is what makes the whole thing possible rather than magical.
At the surface, every problem is a snowflake. Every customer, every outage, every negotiation, every disease is unique in its particulars, and if you stay at the level of particulars you will meet an endless supply of things you have never seen. But the surface is the only layer where that holds. Drop one level, to the relations and constraints and pivots underneath, and the uniqueness collapses. The number of distinct structures is not infinite. It is surprisingly small, and a person tuned to them keeps meeting the same few everywhere, wearing different clothes. Content is unbounded. Structure is finite. That gap between an unbounded surface and a bounded interior is the entire opportunity, and the person who has learned to live at the lower level is fishing in a much smaller pond than everyone working the surface.
This is also why field is a surface-layer idea. Fields are content categories, useful for organizing libraries and universities and résumés, and invisible from where the skill operates. The highway at rush hour and the software release pipeline are different fields. Underneath, they are one shape: flow against a constraint, where the slowest point sets the pace of the whole and the jam arrives before anything is full. One profession meters on-ramps and the other meters work in progress, and neither reads the other’s journals. So when someone carries that shape from one setting to the other, nothing is crossed. The border was drawn through the journals and the job titles, not through the problem. One layer down, the only distinction that exists is the one the whole skill runs on: shapes you know and shapes you do not.
Why we have never been able to see it
If the skill runs one layer below the surface, we have a problem, because nothing we use to sort talent was ever pointed at that layer. The instruments read the surface because the surface is what they were built to read.
The résumé sorts by credential and field, which are content categories. The interview rewards the impressive story, which is a content performance. The nearest miss is the consulting case interview, which was built to watch a candidate structure an ambiguous problem in real time, the cut on demand. It got close enough that an entire coaching industry grew up to defeat it, selling memorized frameworks that let a candidate perform the surface of structuring, and the instrument aimed at the mechanism got captured by the library recital it was designed to bypass. The capture was possible because the case has a key, canonical frameworks and an expected decomposition drawn from one industry’s small stock of scenarios, and anything keyed can be memorized, and anything memorized is content again.
The best prior attempt to name this territory, David Epstein’s Range, got the mechanism right and the cause wrong. Epstein saw the structure layer clearly; the analogical-transfer research is in the book, and he noticed that people who thrive on hard, novel problems often carry a shape in from somewhere unexpected. His conclusion was that breadth builds the capacity: sample widely, accumulate distant analogies, triumph as a generalist. But the research he cites requires something narrower than he claims, variation, not distance. Epstein half-knows it: his own chapter on learning is a tour of varied practice, interleaved problem types building models that travel, and the book’s leap is to scale that finding from mixed problems to mixed careers, further than the research reaches. Schemas form when you meet the same structure under different surfaces, and a deep career in one field supplies that variation by the thousand: no two support calls, no two outages, no two negotiations wear the same clothes. Breadth is one way to vary the surfaces. Depth is the commoner way, and his own best examples confirm it. The designer he admires for building a handheld empire out of a cheap, dated technology was not a generalist wandering between fields; he understood one thing deeply and moved it. The outsiders who crack a field’s problems are experts in their own field first. Range watched the shape travel a spectacular distance and named the capacity after the mileage. But the distance was never the skill. It was the shadow the skill casts on the content layer, and it only becomes visible when the surfaces are far enough apart to look impressive.
That is the quiet catastrophe in how we find talent, and the new vocabulary has not fixed it, only relocated it. We notice the person who solved a biology problem with physics, because the leap is spectacular, and now we also notice the person who performs taste fluently, because the word is fashionable. We do not notice the support rep who read a display fault out of a fuzzy tractor, because the shape landed close to home and nothing looked like genius. Same skill, no spectacle. The distance detectors miss the quiet version; the mystery words can’t detect anything at all, because a word with no mechanism has no instrument. Between the two, the people who do the rare thing all day without looking remarkable stay exactly where they have always been: invisible.
Why it suddenly matters
None of this would be urgent if the world had held still. For most of my career, structure-building was an underrated edge, the kind of thing that let a political science major end up running technology and never quite explain why on paper. It is now the decisive edge, and the thing that changed is the machine.
Artificial intelligence is, at bottom, a content engine, and that includes its impressive moments with structure. Ask it which decision a problem turns on and it will often answer well, because the shapes of many problems have been written down before, and a structure that has been written down is content, retrievable like any other. This is what the dissenters are seeing when they report that the newest models finally show something that feels like judgment, like taste: the AI founder Matt Schumer is right about what he saw, and what he saw was a recorded shape coming back as fluency, because every shape anyone manages to articulate joins the corpus and returns on request. On a problem whose shape has no precedent in its training data, the fluency continues unchanged, the confidence continues unchanged, and nothing in the output marks the line. You can watch this in medicine, where the models have been tested hardest. A team at Shanghai Jiao Tong benchmarked five reasoning models across 1,453 real patient cases, splitting the work by clinical stage. Given a complete workup, the models diagnosed at better than 85% accuracy, and they held that accuracy on rare diseases as well as common ones, because diagnosis is at bottom a match between a presentation and a named condition, and named conditions have written-down criteria. Asked to plan treatment for the same patients, the same models fell to roughly 30%, and here, for the first time, the rare cases did worse than the common ones. The split is not between easy medicine and hard medicine. It is between the step where an answer has already been recorded and the step where a plan has to be built. The machine cannot tell you when it is reciting a known shape and when it is guessing in the same voice, and it cannot reliably catch its own fluent output being structurally wrong. So its structural read is unbankable exactly where the stakes are, on the novel problems, and the only way to know whether its cut is sound is to be able to make the cut yourself.
I ran the old call past a frontier model to see. It took the mapping instantly, tractor means image and fuzzy means resolution, because that inference has been written down over and over since 1999 and a structure that has been written down is content. Then it built. Monitor self-test, cable pins, inline splitters, magnetic interference, the card reseated: each one tested, each one cleared, until the video card was the last component standing and it drafted the warranty claim, ready to send working hardware out for depot repair. The elimination table was clean and every entry in it was true. The conclusion was wrong. The fault was the display driver, which is not an obscure answer. It is the next thing you check after settings, and the model could have retrieved it in a sentence. It never went looking. Each time I closed a door it revised its answer and left the frame alone, and the frame had been hardware since the settings came back clean. The point is not that it missed, because people miss. The point is that nothing in the output marked the miss, and the confidence ran highest at the end. Those are Pirsig’s mechanics, working the wrong build with more conviction at every visit, except that the machine can do it in nine seconds and will do it again for anyone who asks.
This is why the consensus landed where it did. The machine collapsed the cost of producing content while leaving the cost of vouching for it untouched, and when an input turns abundant, advantage migrates to whatever stays scarce beside it. People could feel something human still standing in the flood, and they reached for the nearest words: judgment, for the deciding; taste, for the choosing; agency, for the moving. The words are pointing at the residual. They are not wrong about the location. What they name is the silhouette of the gates seen from the content layer: judgment is what the cut and the break look like from outside, and taste is what a deep library looks like when you cannot see it being consulted. Agency is two words wearing one coat — the direction half is the gates again, and the drive half is not a gate at all but a multiplier on whatever the gates produce.
Where the work is open-ended and the problem’s shape is genuinely new, the gains do not level. They scatter. In the field experiment Fabrizio Dell’Acqua and colleagues ran with 758 BCG consultants, those working inside the machine’s range finished faster and produced better work; those pushed past that range were nineteen percentage points more likely to reach a wrong answer, and seniority did not protect them. Practitioners put the same scatter more bluntly: the same tool that makes one engineer remarkable makes another a net cost to the team. And nobody has yet produced a variable that says in advance which side of that split a person lands on. Tenure does not predict it. Neither does the credential. This essay’s answer is the structure skill, visible now because the machine does the content that used to hide it.
The claim has a boundary, and the best-measured evidence marks it. On bounded work, support queues and routine coding and scripted writing, the machine is not an amplifier but a leveler. Brynjolfsson, Li, and Raymond, following more than five thousand customer-support agents through a rollout, found that the least experienced gained the most while the most experienced gained a little speed and lost a little quality. That is what should happen where the ceiling is a script. What the tool handed over was phrasing and the order of the diagnostic questions, which is content, transferable by definition. The tool also stayed silent whenever it had no training data for the situation, so on the unusual cases there was nothing for the study to measure. Inside the script, the outcome turns on whether you have the content, and the machine hands that to everyone now. Past the script, it turns on structure.
The case I know from the inside
I am a cleaner test of this than most of the people in the talent literature. I am not a generalist, and I have no taste story to sell. The pattern holds anyway.
On paper my path looks incoherent: a political science degree, summers spent farming in Kansas, no computer science, and a career as a chief security officer three times over and now a chief technology officer. But I did not get here by sampling widely. I came up a fairly narrow track, doing what looked from outside like unrelated jobs and felt from the inside like one continuous piece of work. The library started on my grandparents’ farm — the family business my parents had left for office jobs nearby, where my summers and a good share of my weekends went — where a single day hands you an engine, an animal, and the weather, none of it arriving in a state of your choosing, and the verdict cannot be argued with: the engine starts or it does not, the crop comes in or it does not. Farming teaches you early which of the day’s problems will decide the season and which are noise. Political science, absorbed the way I apparently absorbed it, is the study of large-scale systems, where the load sits, which few decisions the whole structure turns on. Technology is where those decisions get made instead of studied, and the wrong ones bring the system down. The résumé could never show the through-line, because a résumé records content, and the through-line was never content. It was one structural read, running under every surface, from a fuzzy tractor forward.
I ran a crude version of this before I had the theory. I made my consultants read NIST 800-27, a security engineering standard long since superseded, because it is the rare NIST document that explains why instead of what and how. I have never stopped doing that. In meetings, in office hours, teaching someone to make a decent cup of coffee, what I am after is the principle underneath, because the principle is the part they can carry into a situation I never trained them for. The ones who took to it are the strongest people I have worked with. The ones who never did stayed on the surface of their jobs for as long as I knew them. That is not a study. It is why I want one built.
What follows, and what we do not yet know
The reform this points to is not the one that arrives every decade under a new name. Teach thinking, not facts is wrong: content is the raw material structure is built from, and a world that trains thinking without content produces confident people with empty libraries. We already train content; that is most of what school and onboarding are. What we leave to accident is everything after it, the cut and the recognition and the break, and the research on transfer says accident is a bad teacher.
In the experiments that founded the transfer field, run by Mary Gick and Keith Holyoak on a problem Karl Duncker had posed decades earlier, people read a story about a general taking a fortress: an army too large for any single mined road, divided into small columns converging from every direction. Minutes later, most of them could not see that the tumor problem in front of them, rays too strong to send down any single path, had the same solution; the shape was in their library and the clothes defeated it. What moved the needle was forcing a comparison, two stories side by side until the shape came loose from its surfaces, and then transfer jumped. Decades of studies since keep finding the same thing: people fail to carry structure forward unless the training forces the rep, shape against shape, across cases that look nothing alike. We even know what the deliberate version looks like at scale, because pockets of it exist and they work, physics classrooms that teach model-building over formula-matching, medical programs that drill the diagnostic pattern rather than the disease list.
The pockets are the indictment. The gates are trainable, the method is known, and no general curriculum makes it the spine. Sometimes the removal was deliberate. Crawford traces what happened to repair manuals: the older ones were written mechanic to mechanic, assuming a reader who held the why and could be trusted with it; once professional technical writers took over, the manuals became pure procedure for a reader presumed to have no judgment at all. The why was not lost. It was edited out. I can add only testimony to the research: nobody ever taught it to me on purpose, and in nearly three decades I have never worked for a company that hired for it, built it, or rewarded it by name. The reform is not less content, and it is not thinking-in-general bolted on beside it. It is content taught as material for abstraction rather than as the destination, the same content delivered in a way that forces the abstraction rep every time, until abstraction becomes a habit rather than a lucky byproduct.
Which returns me to the woman with the fuzzy tractor, and to what I think we are actually missing. The person most likely to thrive beside the machine is not necessarily the one whose brilliance travels a spectacular distance, and not the one who performs the fashionable word most fluently. It is just as likely to be the support rep reading a display fault out of an impossible sentence, doing the rare thing so quietly that no instrument we own would ever have flagged them. Somebody is taking that call right now.
Every instrument we use to choose people measures content, and content is the half the machine now supplies. Our hiring is about to get worse at predicting who will be good at this, not better. Some of the people who can do the other half are already here, unfound, because nothing we own looks one layer down. And the capacity can be taught: the research has known it for decades, the pockets that do it work, and no curriculum has made it the spine. Build the instrument that can see one layer down. Build the curriculum that makes more of them. The instrument finds the ones accident already made; the curriculum stops trusting accident with the making. We are short of these people. We don’t have to be.
Sources
The consensus. Paul Graham, X, 14 February 2026: in the AI age taste becomes more important, and when anyone can make anything the big differentiator is what you choose to make. He linked his 2002 essay “Taste for Makers.” Two days later, on 16 February, OpenAI president and co-founder Greg Brockman posted that taste is a new core skill. Both verified against the original posts and contemporaneous coverage; note that the Brockman line is widely misattributed to Sam Altman.
The dissent. Matt Schumer (co-founder and CEO, OthersideAI), February 2026 viral essay and subsequent X posts: OpenAI’s GPT-5.3 Codex model showed “something that felt, for the first time, like judgment. Like taste,” adding on X that he sees nothing about taste or direction as uniquely human, since anything a model can train on it can learn. Reported in Fortune, 27 February 2026.
The tests that carry the name. Situational judgment tests are validated predictors of job performance (McDaniel et al., meta-analyses 2001 and 2007; Sackett et al. 2022), but the literature’s own long-running problem is that SJTs are a measurement method rather than an identified construct: they score whether a respondent selects the answer subject-matter experts keyed as effective, and they correlate heavily with general cognitive ability. The essay’s claim is about what no instrument reaches, not about SJT validity.
Transfer and the fortress/tumor pair. Karl Duncker, On Problem-Solving (1945), for the radiation problem. Mary Gick and Keith Holyoak, “Analogical Problem Solving,” Cognitive Psychology 12 (1980), and “Schema Induction and Analogical Transfer,” Cognitive Psychology 15 (1983) — spontaneous transfer from the military story to the radiation problem was rare without a hint; forced comparison across two analogues induced the schema and raised transfer sharply.
Depth, manuals, and the tappets. Matthew B. Crawford, Shop Class as Soulcraft: An Inquiry Into the Value of Work (Penguin, 2009). The motorcycle story Crawford retells is from Robert M. Pirsig, Zen and the Art of Motorcycle Maintenance (1974).
Range. David Epstein, Range: Why Generalists Triumph in a Specialized World (Riverhead, 2019).
The leveler on bounded work. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140:2 (May 2025), 889–942. 5,172 customer-support agents; productivity up 15% on average; less experienced and lower-skilled workers improved in speed and quality, while the most experienced and highest-skilled saw small gains in speed and small declines in quality. https://doi.org/10.1093/qje/qjae044
The confident wrong answer. Fabrizio Dell’Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality.” 758 BCG consultants; inside the frontier, 12.2% more tasks completed, 25.1% faster, 40% higher quality; outside it, 19 percentage points more likely to reach a wrong answer. Harvard Business School working paper 24-013 (2023), subsequently published in Organization Science.
Recorded answers versus built plans. Pengcheng Qiu, Chaoyi Wu, et al., “Quantifying the reasoning abilities of LLMs on clinical cases,” Nature Communications (2025). MedR-Bench: 1,453 structured patient cases across 13 body systems and 10 specialties, 656 of them rare-disease cases, evaluated across examination recommendation, diagnosis, and treatment planning. In the oracle diagnostic setting every model exceeded 80% accuracy and rare-disease performance held (DeepSeek-R1: 89.8% overall, 91.0% on rare cases). Treatment-planning accuracy for the same models ran near 30%, and the paper notes that unlike diagnosis, where rare cases do not impact performance, treatment planning declines further for rare diseases. https://www.nature.com/articles/s41467-025-64769-1