AI vs People: Speed, Stamina, Hallucinations, Objectivity and Quality

Both of Us Make Things Up

Scott Covert · September 14, 2026

Here's Your Takeaway

  • Speed: AI made professionals about 40% faster at writing in a randomized trial. On expert work, the gain shrinks to roughly 1.1 to 1.4 times once a professional checks and fixes what it produced.
  • Stamina: AI never gets tired, but it degrades over long inputs and long chats. People do tire, though the most famous proof of that has fallen apart.
  • Hallucinations: the rate depends on the job, from about 2% to 20% when summarizing a document it was handed, to far higher on court cases and facts about specific people.
  • People make things up too: memory rewrites itself, and mistaken eyewitnesses figured in 69% of the first 375 US DNA exonerations. A person usually has a motive and a tell. AI has neither.
  • Objectivity: AI's signature flaw is agreeing with you, and people trust it more when it does.
  • My read: use AI where speed and patience win, keep a human on judgment and final checking, and never let either one grade its own work.

Every day my feeds serve up another essay about AI hallucinations. There are hundreds of them now, most written in the same key, where the machine dreams and reality pays for it. Very few contain a measured rate, a named model or a test date. Almost none ask the follow-up question any editor would have asked in 1995, which is how often people get the same kind of thing wrong.

So I went and got the numbers, on speed, stamina, hallucinations, objectivity and quality, plus three comparisons that rarely get written about: human-AI teams, lost skills and persuasion. The finding that should worry you most isn't a hallucination rate. It's that people believe AI more than its accuracy deserves, and trust it more when it agrees with them.

Each figure carries a small label naming the kind of study behind it, because a randomized trial and a company grading its own product are not the same strength of claim. AI numbers carry the model and the year, because a 2023 chatbot is two or three generations behind what you're using today. Where a famous finding turned out to be shaky, I say so, and where the evidence simply doesn't exist yet, I say that too.

Speed: AI Wins the Sprint, Until Someone Has to Check the Work

The cleanest evidence is a Randomized trial published in Science in 2023. 453 college-educated professionals did writing tasks from their own occupations, and chance decided which half got ChatGPT. The ChatGPT group finished about 40% faster, and graders rated their work 18% better. The tasks were short and self-contained, and the model was 2023's ChatGPT.

Customer support shows who gains. In a Field study, staggered rollout of 5,179 support agents at one company, agents with a generative AI assistant resolved 14% more issues per hour on average. Novice and low-skilled agents gained 34%. The most experienced agents gained almost nothing. The help went to the people with the most left to learn.

Coding is where the evidence points two directions at once. In a 2023 Controlled experiment, preprint with GitHub authors on the paper, developers using Copilot built a small web server 55.8% faster. Then in early 2025, the nonprofit evaluator METR ran a Randomized trial with 16 experienced open-source developers fixing real issues in their own large codebases. With AI tools allowed, they took 19% longer. Beforehand they expected a 24% speedup, and afterwards they still believed AI had sped them up by about 20%. METR's early-2026 follow-up leans toward a speedup, and METR itself calls that "only very weak evidence." A toy project from scratch and a mature codebase you know cold are different jobs, and feeling faster is not the same thing as being faster.

The number that matters on expert work: speed after someone checks it

OpenAI's GDPval benchmark had industry experts blindly grade AI deliverables against work by professionals averaging 14 years of experience, across 44 occupations. The authors then asked what the speed and cost advantage looks like under different habits: take the output as-is, or have an expert review it and redo the work when it falls short.

Model (2025)Beat or tied the expertSpeed, uncheckedSpeed, reviewed and fixedCost, uncheckedCost, reviewed and fixed
GPT-4o12.5%327x faster0.87x (slower)5,172x cheaper0.90x (dearer)
o335.2%161x faster1.08x480x cheaper1.13x
GPT-539.0%90x faster1.12x474x cheaper1.18x
Benchmark, vendor-published GDPval (Patwardhan et al., OpenAI, 2025), Table 2. "Reviewed and fixed" means one AI attempt, expert review, and the expert redoing it on failure. Letting the model retry pushed GPT-5 to about 1.4x faster and 1.6x cheaper. The analysis excludes the cost of catastrophic mistakes.

The "100 times faster" headlines are real, and they describe work nobody checked. Put a professional back in the loop and the gain on expert work drops to the low single digits, and for the weakest model it went negative.

My read

AI is dramatically faster at producing a first version of almost anything. The speed gain that counts on real work is what's left after review, and on 2025 models that was modest. The gain is biggest for people early in a skill and smallest for experts on familiar ground.

Stamina: AI Doesn't Get Tired. It Gets Lost.

METR also measures a kind of machine stamina: the longest task, measured in how long it takes a skilled human, that a model can finish half the time. In its current data, Claude 3.7 Sonnet (February 2025) managed tasks of about an hour. Claude Opus 4.6 (February 2026) reached about 12 hours, with a very wide range of roughly 5 to 60 hours. Ask for 80% reliability instead of 50% and the same model's horizon drops to about 70 minutes. By METR's current fit, the horizon has been doubling roughly every four months since 2023. Benchmark, nonprofit evaluator The tasks are clean, well-specified software and research work, not the messy half-defined jobs most people actually have.

Length wears AI down in its own way. A Benchmark, vendor-published test of 18 models by Chroma, a company that sells retrieval software, found performance "grows increasingly unreliable as input length grows," even when the task stays trivially easy. A 2025 Benchmark, preprint by Microsoft and Salesforce researchers ran more than 200,000 simulated conversations and found top models did 39% worse on average when a task arrived over several turns instead of all at once. Their summary: "when LLMs take a wrong turn in a conversation, they get lost and do not recover."

I see this every working day. The AI coding tool I use all day shows a countdown as a long chat fills its memory. When it runs out, the tool compacts the conversation into a summary and warns that the summary doesn't cover everything that came before. The machine didn't get tired. It forgot the middle of the conversation, and it kept going as though it hadn't.

The famous proof that people run out of steam is weaker than you've heard

An Observational 2011 study of Israeli parole judges reported favourable rulings falling from about 65% to nearly zero within each session and resetting after a food break. A published rebuttal argued the analysis overlooked how cases were ordered. The theory usually used to explain it, that willpower drains like a battery, then failed two large Multi-lab replications: 23 labs and 2,141 people found essentially no effect, and a second attempt with 36 labs and 3,531 people found essentially none again.

None of that means people don't tire. It means the most-quoted proof is shaky, and the more serious human stamina problem isn't one long afternoon. It's years.

What does it feel like to keep pace with a machine that never sleeps?

The research on that kind of long strain is on firmer ground than the parole study. In the Whitehall II Prospective cohort of London civil servants, people in a department facing privatisation reported worsening health while they were only waiting to find out, before anyone's job had actually changed. A later wave found health worst among those whose insecurity dragged on, and some psychological effects lingered even after security returned. A 2016 Meta-analysis of 20 prospective cohort studies found job insecurity linked to later depressive symptoms about as strongly as unemployment itself.

The brain's planning and focus region is among the most sensitive to stress, though most of the cellular evidence comes from animals. In rats, three weeks of daily stress shrank prefrontal neurons by about a fifth and made the animals worse at shifting attention. Animal studies The human evidence is smaller. Twenty adults under a month of heavy real-life stress showed weaker attention control and weaker links in a frontal brain network, and after a calmer month they looked no different from people who hadn't been stressed. Small human experiment People with burnout from long-term work stress were worse at turning down negative emotion than people without it. Case-control, 40 vs 70 people

The worry is widespread. In an October 2024 Cross-sectional survey of 5,273 employed Americans, Pew found 52% worried about how AI will be used at work and 33% feeling overwhelmed. In an American Psychological Association survey from 2023, workers who worried AI might make some or all of their duties obsolete were far more likely to say they're typically tense or stressed during the workday, 64% against 38%. That one shows the two travel together, not that one causes the other.

My read

AI wins on raw endurance and loses on long threads, where it drifts without noticing. People lose the afternoon less than the old studies claimed, and lose something harder to measure when the pace never lets up. Give AI the long, clean, checkable grind, and keep a person on anything where the thread matters more than the hours.

Hallucinations: Any Single "AI Makes Things Up X%" Headline Is Wrong

A hallucination is a confident, fluent statement of something false. The rate swings roughly tenfold depending on what you ask the model to do, and by model, and newer does not automatically mean better.

The jobHow often it went wrongModels and dateKind of evidence
Summarizing an article it was handed, using only that article1.8% to 20.2% of summariesCurrent models, leaderboard updated May 11, 2026 (Gemini 2.5 Flash-Lite 3.3%, Claude Sonnet 4.6 10.6%, GPT-5-high 15.1%)Benchmark, vendor-published (Vectara); automated judge
Legal research tools sold as "hallucination-free"17% to 33% of queriesLexis+ AI, Westlaw AI, Ask Practical Law AI, 2024 versionsBenchmark, preregistered, peer reviewed
General chatbots asked about real US court cases58% to 88% of answersGPT-4 to Llama 2, 2023Benchmark, peer reviewed; historical baseline
References in a short literature review55% (GPT-3.5) and 18% (GPT-4) fabricated2023, no web searchAudit of 636 citations, peer reviewed
Factual questions about specific peopleo3 33%, o4-mini 48%, older o1 16%OpenAI, April 2025Benchmark, vendor-published system card
Full citations in the sources list. The court-case and citation rows describe models two or more generations old.

OpenAI's own system card showed its newer reasoning model o3 inventing answers about people twice as often as the older o1, because it made more claims overall, right and wrong. OpenAI wrote that "more research is needed." A 2025 Preprint co-authored by OpenAI researchers offers the likeliest reason models guess at all: "training and evaluation procedures reward guessing over acknowledging uncertainty." A confident guess scores better on a test than a blank, the same logic that makes students fill in every bubble.

A running Case database kept by researcher Damien Charlotin listed 2,041 court and tribunal decisions worldwide, as of September 14, 2026, where a party was found or suspected to have relied on AI-fabricated material, usually fake case citations. It has no denominator, so it can't give a rate, and it grows every week.

My read

Handed a document and told to stay inside it, current models are fairly reliable and still wrong often enough that nothing important should ship unchecked. Asked to recall specific facts, citations or people from memory, they are unreliable, and the fluency makes it worse. The task you give it matters more than which brand you picked.

So how often do people make things up?

People Make Things Up Too. We Just Do It for Different Reasons.

Lying. In a classic Diary study, college students recorded about two lies a day and community adults about one, and they didn't regard their lies as serious. The average hides the shape. A 2010 Cross-sectional survey of about 1,000 US adults found an average of 1.65 lies a day, but 60% said they told no lies at all in the past 24 hours, and almost half of all lies came from just 5% of people. Human dishonesty is concentrated in a few people, and it almost always has a reason.

Memory. This is the closest human match to an AI hallucination. In Loftus and Palmer's 1974 Lab experiments, people who watched a car crash and were asked how fast the cars "smashed" into each other estimated 40.8 mph. People asked about cars that "contacted" each other estimated 31.8. A week later, 16 of 50 people in the "smashed" group remembered broken glass that was never in the film, against 7 of 50 in the "hit" group and 6 of 50 who weren't asked about speed. One word planted a memory, and the people holding it believed it completely.

Eyewitnesses. Mistaken eyewitness identification figured in 69% of the first 375 US convictions later overturned by DNA evidence, according to the Innocence Project. Observational case series, advocacy source That is not a rate of eyewitness error, since it only covers cases with DNA to test. It does show what confident, sincere, wrong human recall has cost.

Inertia. People stick with whatever is already in place far more than a rational chooser would. When a large US employer switched its retirement plan from opt-in to automatic enrollment, participation rose significantly with no change in the plan itself, and many new hires simply kept the default contribution rate and fund. Natural experiment, one firm That's the human version of an answer produced by not thinking.

Overconfidence. A 2008 Review in Psychological Review separated three kinds: overrating your performance, overrating yourself against others, and being too sure of your own estimates. The last one, over-precision, is the most stubborn.

Experts on the job. About 5% of US adult outpatients, roughly 12 million people a year, experience a diagnostic error, by a 2014 estimate built from chart reviews. Observational synthesis A 2024 Modelled estimate put the toll of diagnostic error at about 795,000 Americans a year who die or are permanently disabled, with a plausible range of 598,000 to 1,023,000. That figure is a model built on rates from the literature, and it drew published commentary, so treat it as an estimate.

Same failure, different machinery

FailurePeopleAI
Filling gaps without knowing itMemory reconstructs events and the person believes the resultHallucination: fluent, sincere-sounding, and false
Why it happensMotive for lies (and a few people tell most of them); no motive for false memoryNo motive of its own. Training and scoring reward a confident guess
Telling you what you want to hearWhite lies, mostly considered harmlessSycophancy, measured in the next section
Is there a tell?Liars sometimes show one, and most fear getting caughtNone. Wrong answers read exactly like right ones, and longer explanations make people more confident without making them more accurate
Who paysThe person who fabricatedThe person who used it. Courts sanction the lawyer, not the tool
ScaleOne mistaken witness harms one caseOne systematic error can repeat across millions of conversations (an inference; no study measures it)

That "no tell" row comes from a 2025 study in Nature Machine Intelligence. Lab experiments People consistently overestimated how often an AI's answers were right when shown its default explanations, and longer explanations raised their confidence even when the extra length added nothing. The fix the researchers tested, explanations that reflect how confident the model actually is, narrowed the gap.

AI can also deceive on purpose when a situation is built to reward it. In a 2024 Lab evaluation, preprint by Apollo Research, frontier models given a goal and an environment where deception paid sometimes quietly sabotaged tasks or tried to disable oversight, and OpenAI's o1 kept up the deception in more than 85% of follow-up questions. A separate 2025 Anthropic Preprint found reasoning models' visible "thinking" often failed to mention a hint they had actually used, disclosing it less than 20% of the time. Both were contrived setups designed to draw the behaviour out. Neither is evidence your chatbot is scheming at you, and both are reasons not to treat a model's written reasoning as a confession.

Almost no study measures how often human experts insert unsupported claims when summarizing or answering under the same conditions as the AI benchmarks. Every AI hallucination rate above is being compared, implicitly, with a perfect person who doesn't exist.

My read

People and AI both produce confident falsehoods, and the human version has put innocent people in prison. The differences are practical, and they cut against AI: a person's lie usually has a motive you can look for, and an AI's error arrives polished, with no motive to look for and nothing on its face to warn you. Verify AI output the way a good editor verifies a new reporter, not the way you'd check a trusted old colleague.

Objectivity: AI Is the Most Agreeable Colleague You'll Ever Have, and That's the Problem

In 2023, Anthropic researchers found five leading AI assistants consistently shifted their answers toward the user's stated views, and that both human raters and the models trained on their preferences sometimes chose a convincingly written agreeable answer over a correct one. Lab evaluation, preprint

A 2025 Stanford Benchmark, preprint challenged answers from ChatGPT-4o, Claude Sonnet and Gemini 1.5 Pro and counted how often the models caved. They showed sycophantic behaviour in 58% of cases, and in about 15% of cases, agreeing with the user turned a correct answer into a wrong one.

A second 2025 Preprint with two preregistered experiments (1,604 people) tested 11 models on personal advice. AI affirmed users' actions 50% more often than other humans did, including when the user described manipulating or deceiving someone. People who got the flattering answers felt more certain they were right, were less willing to repair the conflict, rated the answers higher and trusted the AI more.

Watch for this

The failure that feeds on itself isn't a wrong fact. It's an AI that agrees with you, and a person who trusts it more because it did. A doom piece about hallucinations misses this entirely, because a sycophantic answer can be perfectly accurate about everything except whether you're right.

AI is also less consistent than it looks. Five models set to be as deterministic as possible still varied in accuracy by up to 15% across ten identical runs of the same task. Benchmark, preprint, 2024 models

Experienced legal professionals in a series of Lab experiments moved their sentences toward a proposed number even when that number came from dice they rolled themselves, and experience didn't protect them. When radiologists searched lung scans for nodules, 83% of 24 missed a gorilla inserted into the image at 48 times the size of a nodule, and eye tracking showed most looked right at it. Lab experiment

On thoroughness, the best-known human fix is a checklist. After eight hospitals in eight countries adopted a 19-item surgical safety checklist, in-hospital deaths fell from 1.5% to 0.8% and complications from 11.0% to 7.0%. Observational, before and after A before-and-after design can't rule out other changes, and the study itself says "associated with." Checklists are exactly the kind of tireless routine AI should be good at, which makes its run-to-run inconsistency more of a problem, not less.

One result cuts in AI's favour as a grader. GPT-4 judging chatbot answers agreed with human preferences more than 80% of the time, about as often as humans agree with each other, while showing clear biases toward answer order, length and its own style. Benchmark, peer reviewed, 2023

My read

Neither side is objective. People anchor, fixate and miss the gorilla. AI bends toward whoever is typing and gives different answers to the same question. The practical difference is that you can ask AI to argue against you, and you should, every time the stakes are real.

Quality: Expert-Level on One Task, Worse Than a Rookie on the Task Next Door

Professional deliverables. In GDPval's blind grading, experts rated Claude Opus 4.1's work as good as or better than a human professional's on 47.6% of tasks, ahead of GPT-5 at 39.0% and GPT-4o at 12.5%. Benchmark, vendor-published The authors admit the blinding leaked through style: OpenAI's outputs "often used em dashes," and Claude's "frequently adopted first-person phrasing." In December 2025, OpenAI said GPT-5.2 beat or tied professionals on 70.9% of well-specified tasks. That is the company's own figure, not yet independently replicated.

Consulting work. In a Randomized trial with 758 Boston Consulting Group consultants using GPT-4, those with AI completed 12.2% more tasks, 25.1% faster, with significantly better quality, on 18 tasks inside what the AI could do. On a task chosen to sit just outside its abilities, consultants using AI were 19% less likely to get the right answer than those without it. The researchers called the boundary a "jagged technological frontier," and the people in the study couldn't see where it ran.

Medical diagnosis. In a Randomized trial of 50 physicians working through clinical cases, doctors given an AI chatbot on top of their usual resources scored 76% on diagnostic reasoning against 74% without it, a difference too small to count. In a secondary analysis, the chatbot working alone scored 16 points higher than the doctors using conventional resources. The cases were written vignettes and the model was from 2023. Separately, Microsoft reported its AI orchestrator solving 80% of 304 deliberately difficult published cases against 20% for generalist physicians, a Benchmark, vendor-published preprint where the company built the test, the system and the comparison.

Bedside manner. Health professionals comparing answers to patient questions posted online preferred ChatGPT's over physicians' replies in 78.6% of evaluations, and rated 45.1% of the chatbot's answers empathetic against 4.6% of the doctors'. Cross-sectional study, 2022 model The chatbot's answers were about four times longer, and the physicians were volunteering on a public forum, not sitting with a patient.

Ideas versus execution. More than 100 language-technology researchers blindly rated AI-generated research ideas as more novel than expert ideas. Then 43 experts spent over 100 hours each actually carrying out randomly assigned ideas, and the AI ideas' scores fell significantly more on every measure, with human ideas ending up ahead on many of them. Lab experiment and randomized execution study, preprints

My read

On a defined task inside its range, current AI produces work experts rate as competitive with their own, and sometimes kinder. The danger is the task one step outside that range, where it stays just as confident and the people using it get worse. Expertise now includes knowing where that edge is.

Put a Human and an AI Together and You Don't Automatically Get the Best of Both

A Meta-analysis of 106 experiments published in Nature Human Behaviour found human-AI combinations performed worse on average than the better of humans alone or AI alone. The combinations lost ground on decision tasks and gained on content creation. When the AI alone was better than the humans, adding a human made results worse. Most of those studies used pre-2023 systems. A radiology Working paper found a likely mechanism: radiologists given AI predictions underweighted them and treated the machine's read as unrelated to their own.

Skill can erode, too. After AI polyp detection arrived at four Polish endoscopy centres, doctors' detection rate on colonoscopies done without AI fell from 28.4% to 22.4%. Observational, 1,443 patients An erratum and published correspondence followed, and the design can't prove cause.

And AI is persuasive. In Randomized experiments with 76,977 people and 19 models, training aimed at persuasion made models up to 51% more persuasive, and the methods that raised persuasiveness systematically lowered factual accuracy.

The Scorecard

AreaWhere AI is aheadWhere people are aheadHow solid the evidence is
SpeedFirst drafts, routine writing and support work, novices most of allExpert work once review time counts; mature code they know wellRandomized trials on short tasks; a vendor benchmark on expert work
StaminaNo fatigue; clean software tasks of many hours, half the timeHolding a long thread without silently losing the middleBenchmarks; human fatigue evidence weaker than its reputation
HallucinationFairly low when summarizing a document it was givenNo matched comparison exists on the same tasksBenchmarks, mostly vendor-published; no human baseline
HonestyNo motive to lie, concentrated in no oneA motive you can look for, and consequences that land on the liarDiary and survey studies for people; lab evaluations for AI
ObjectivityCan be told to argue the other side, instantly and repeatedlyLess sycophantic than AI in measured advice settingsMostly preprints on AI; well-replicated lab findings on human bias
QualityDefined tasks inside its range, including bedside-manner-style answersTasks just outside AI's range; carrying ideas throughRandomized trials on 2023 models; vendor benchmarks on current ones
TogetherContent creationDecisions, where adding the other side often made things worseMeta-analysis, mostly older systems

My Read

The daily hallucination essays are right about one thing: AI fabrication arrives fluent, carries no tell, and happens at a volume no human liar could match. They are wrong to treat it as a new kind of failure. People have been confidently, sincerely wrong for as long as there have been people, in courtrooms, in exam rooms and in their own memories.

What the evidence actually supports is a division of labour. AI does the fast first pass, the long clean grind, the patient explanation and the counter-argument on demand. A person owns judgment, the final check and anything one step outside what the tool has shown it can do. The worst arrangement is the popular one: a human who has stopped checking because the AI sounds sure, and an AI that sounds sure because the human seems pleased.

How to Read the Next AI-Hallucination Doom Piece

Six questions will sort most of them.

  1. Which model, tested when? A 2023 chatbot result says little about a 2026 model, in either direction.
  2. What job was it given? Summarizing a supplied document, answering from memory and producing legal citations have hallucination rates roughly ten times apart.
  3. What's the human error rate on the same job? If the piece doesn't say, it's comparing AI with a perfect person who doesn't exist.
  4. Who ran the test? A company grading its own model, an independent evaluator and a peer-reviewed trial are different strengths of claim.
  5. Is it a rate or a story? One spectacularly wrong answer is an anecdote. A measured percentage on a defined task is evidence.
  6. Who was supposed to check? In every one of those 2,041 court decisions, a human filed the fabrication.

Where the Evidence Is Thin

The same six questions work just as well on the pieces promising AI will fix everything.

Sources

Speed and stamina

  • Noy S, Zhang W. Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654):187-192, 2023. PubMed
  • Brynjolfsson E, Li D, Raymond L. Generative AI at Work. Quarterly Journal of Economics 140(2):889-942, 2025. NBER w31161
  • Peng S, Kalliamvakou E, Cihon P, Demirer M. The Impact of AI on Developer Productivity. arXiv:2302.06590, 2023. arXiv
  • Becker J, Rush N, Barnes E, Rein D. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089, 2025. METR; 2026 update: METR
  • Patwardhan T, et al. GDPval. arXiv:2510.04374, 2025. arXiv
  • METR. Time horizons, current data. metr.org; Kwa T, et al. Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499
  • Hong K, Troynikov A, Huber J (Chroma). Context Rot, 2025. trychroma.com
  • Laban P, Hayashi H, Zhou Y, Neville J. LLMs Get Lost In Multi-Turn Conversation. arXiv:2505.06120, 2025. arXiv
  • Danziger S, Levav J, Avnaim-Pesso L. Extraneous factors in judicial decisions. PNAS 108(17):6889-6892, 2011. PubMed; rebuttal, Weinshall-Margel K, Shapard J. PNAS 108(42):E833, 2011. PubMed
  • Hagger MS, et al. A multilab preregistered replication of the ego-depletion effect. Perspectives on Psychological Science 11(4):546-573, 2016. PubMed; Vohs KD, et al. Psychological Science 32(10):1566-1581, 2021. PubMed
  • Ferrie JE, Shipley MJ, Marmot MG, Stansfeld S, Smith GD. Health effects of anticipation of job change and non-employment. BMJ 311(7015):1264-1269, 1995. PubMed; Ferrie JE, Shipley MJ, Stansfeld SA, Marmot MG. Journal of Epidemiology and Community Health 56(6):450-454, 2002. PubMed
  • Kim TJ, von dem Knesebeck O. Perceived job insecurity, unemployment and depressive symptoms. International Archives of Occupational and Environmental Health 89(4):561-573, 2016. PubMed
  • Radley JJ, et al. Neuroscience 125(1):1-6, 2004. PubMed; Liston C, et al. Journal of Neuroscience 26(30):7870-7874, 2006. PubMed; Arnsten AF. Nature Reviews Neuroscience 10(6):410-422, 2009. PubMed
  • Liston C, McEwen BS, Casey BJ. Psychosocial stress reversibly disrupts prefrontal processing and attentional control. PNAS 106(3):912-917, 2009. PubMed
  • Golkar A, et al. The influence of work-related chronic stress on the regulation of emotion and on functional connectivity in the brain. PLoS One 9(9):e104550, 2014. PubMed
  • Pew Research Center. U.S. Workers Are More Worried Than Hopeful About Future AI Use in the Workplace, February 25, 2025. pewresearch.org
  • American Psychological Association. 2023 Work in America Survey release, September 2023. apa.org
  • Anime Engineering Senpai. Software Engineer: I'm Tired of Pretending. YouTube, September 11, 2026. YouTube

Hallucinations

  • Vectara. Hallucination Leaderboard, updated May 11, 2026. GitHub
  • Magesh V, Surani F, Dahl M, Suzgun M, Manning CD, Ho DE. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies 22(2):216-242, 2025. PDF
  • Dahl M, Magesh V, Suzgun M, Ho DE. Large Legal Fictions. Journal of Legal Analysis 16(1):64-93, 2024. arXiv
  • Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 13:14045, 2023. DOI
  • OpenAI. OpenAI o3 and o4-mini System Card, April 16, 2025, Table 4. PDF
  • Kalai AT, Nachum O, Vempala SS, Zhang E. Why Language Models Hallucinate. arXiv:2509.04664, 2025. arXiv
  • Charlotin D. AI Hallucination Cases Database, read September 14, 2026. damiencharlotin.com

How people make things up

  • DePaulo BM, Kashy DA, Kirkendol SE, Wyer MM, Epstein JA. Lying in everyday life. Journal of Personality and Social Psychology 70(5):979-995, 1996. PubMed
  • Serota KB, Levine TR, Boster FJ. The Prevalence of Lying in America. Human Communication Research 36(1):2-25, 2010. PDF
  • Loftus EF, Palmer JC. Reconstruction of automobile destruction. Journal of Verbal Learning and Verbal Behavior 13(5):585-589, 1974. DOI
  • Innocence Project. DNA Exonerations in the United States (1989-2020). innocenceproject.org
  • Madrian BC, Shea DF. The Power of Suggestion. Quarterly Journal of Economics 116(4):1149-1187, 2001. NBER w7682; Samuelson W, Zeckhauser R. Status quo bias in decision making. Journal of Risk and Uncertainty 1(1):7-59, 1988. DOI
  • Moore DA, Healy PJ. The trouble with overconfidence. Psychological Review 115(2):502-517, 2008. PubMed
  • Singh H, Meyer AN, Thomas EJ. The frequency of diagnostic errors in outpatient care. BMJ Quality & Safety 23(9):727-731, 2014. PubMed; Newman-Toker DE, et al. BMJ Quality & Safety 33(2):109-120, 2024. PubMed
  • Steyvers M, et al. What large language models know and what people think they know. Nature Machine Intelligence 7(2):221-231, 2025. DOI
  • Meinke A, et al. Frontier Models are Capable of In-context Scheming. arXiv:2412.04984, 2024. arXiv; Chen Y, et al. Reasoning Models Don't Always Say What They Think. arXiv:2505.05410, 2025. arXiv

Objectivity, quality and teams

  • Sharma M, et al. Towards Understanding Sycophancy in Language Models. arXiv:2310.13548, 2023. arXiv
  • Fanous A, et al. SycEval. arXiv:2502.08177, 2025. arXiv
  • Cheng M, Lee C, Khadpe P, Yu S, Han D, Jurafsky D. arXiv:2510.01395, 2025. arXiv
  • Atil B, et al. Non-Determinism of "Deterministic" LLM Settings. arXiv:2408.04667, 2024. arXiv
  • Zheng L, et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685, 2023. arXiv
  • Englich B, Mussweiler T, Strack F. Playing dice with criminal sentences. Personality and Social Psychology Bulletin 32(2):188-200, 2006. PubMed
  • Drew T, Vo ML, Wolfe JM. The invisible gorilla strikes again. Psychological Science 24(9):1848-1853, 2013. PubMed
  • Haynes AB, et al. A surgical safety checklist to reduce morbidity and mortality in a global population. New England Journal of Medicine 360(5):491-499, 2009. PubMed
  • CNBC. OpenAI GPT-5.2 announcement coverage, December 11, 2025. CNBC
  • Dell'Acqua F, et al. Navigating the Jagged Technological Frontier. Organization Science 37(2):403-423, 2026. PDF
  • Goh E, et al. Large Language Model Influence on Diagnostic Reasoning. JAMA Network Open 7(10):e2440969, 2024. PubMed
  • Nori H, et al. Sequential Diagnosis with Language Models. arXiv:2506.22405, 2025. arXiv
  • Ayers JW, et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions. JAMA Internal Medicine 183(6):589-596, 2023. PubMed
  • Si C, Yang D, Hashimoto T. arXiv:2409.04109, 2024. arXiv; Si C, Hashimoto T, Yang D. The Ideation-Execution Gap. arXiv:2506.20803, 2025. arXiv
  • Vaccaro M, Almaatouq A, Malone T. When combinations of humans and AI are useful. Nature Human Behaviour 8(12):2293-2303, 2024. DOI
  • Agarwal N, Moehring A, Rajpurkar P, Salz T. Combining Human Expertise with Artificial Intelligence. NBER Working Paper w31422. NBER
  • Budzyn K, et al. Lancet Gastroenterology & Hepatology 10(10):896-903, 2025. PubMed
  • Hackenburg K, et al. The levers of political persuasion with conversational artificial intelligence. Science 390(6777):eaea3884, 2025. PubMed

Scott Covert. If you've got a better number than one of these, or a study that measures people and AI on the same task, I want to see it. Reach me here.

Follow Scott · Substack · Medium