Close Menu
New York Examiner News

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Blink-182’s Mark Hoppus reveals how he sang on unreleased Linkin Park song

    September 27, 2026

    Coastal erosion is so bad in California that this mostly GOP city may tax itself to restore beach

    September 27, 2026

    Trump Can’t Get Even Go To The Golf Course Without Being Called A Pedo

    September 27, 2026
    Facebook X (Twitter) Instagram
    New York Examiner News
    • Home
    • US News
    • Politics
    • Business
    • Science
    • Technology
    • Lifestyle
    • Music
    • Television
    • Film
    • Books
    • Contact
      • About
      • Amazon Disclaimer
      • DMCA / Copyrights Disclaimer
      • Terms and Conditions
      • Privacy Policy
    New York Examiner News
    Home»Science»AI hallucinations are getting worse – and they’re here to stay
    Science

    AI hallucinations are getting worse – and they’re here to stay

    By AdminMay 12, 2025
    Facebook Twitter Pinterest LinkedIn WhatsApp Email Reddit Telegram
    AI hallucinations are getting worse – and they’re here to stay


    AI hallucinations are getting worse – and they’re here to stay

    Errors tend to crop up in AI-generated content

    Paul Taylor/Getty Images

    AI chatbots from tech companies such as OpenAI and Google have been getting so-called reasoning upgrades over the past months – ideally to make them better at giving us answers we can trust, but recent testing suggests they are sometimes doing worse than previous models. The errors made by chatbots, known as “hallucinations”, have been a problem from the start, and it is becoming clear we may never get rid of them.

    Hallucination is a blanket term for certain kinds of mistakes made by the large language models (LLMs) that power systems like OpenAI’s ChatGPT or Google’s Gemini. It is best known as a description of the way they sometimes present false information as true. But it can also refer to an AI-generated answer that is factually accurate, but not actually relevant to the question it was asked, or fails to follow instructions in some other way.

    An OpenAI technical report evaluating its latest LLMs showed that its o3 and o4-mini models, which were released in April, had significantly higher hallucination rates than the company’s previous o1 model that came out in late 2024. For example, when summarising publicly available facts about people, o3 hallucinated 33 per cent of the time while o4-mini did so 48 per cent of the time. In comparison, o1 had a hallucination rate of 16 per cent.

    The problem isn’t limited to OpenAI. One popular leaderboard from the company Vectara that assesses hallucination rates indicates some “reasoning” models – including the DeepSeek-R1 model from developer DeepSeek – saw double-digit rises in hallucination rates compared with previous models from their developers. This type of model goes through multiple steps to demonstrate a line of reasoning before responding.

    OpenAI says the reasoning process isn’t to blame. “Hallucinations are not inherently more prevalent in reasoning models, though we are actively working to reduce the higher rates of hallucination we saw in o3 and o4-mini,” says an OpenAI spokesperson. “We’ll continue our research on hallucinations across all models to improve accuracy and reliability.”

    Some potential applications for LLMs could be derailed by hallucination. A model that consistently states falsehoods and requires fact-checking won’t be a helpful research assistant; a paralegal-bot that cites imaginary cases will get lawyers into trouble; a customer service agent that claims outdated policies are still active will create headaches for the company.

    However, AI companies initially claimed that this problem would clear up over time. Indeed, after they were first launched, models tended to hallucinate less with each update. But the high hallucination rates of recent versions are complicating that narrative – whether or not reasoning is at fault.

    Vectara’s leaderboard ranks models based on their factual consistency in summarising documents they are given. This showed that “hallucination rates are almost the same for reasoning versus non-reasoning models”, at least for systems from OpenAI and Google, says Forrest Sheng Bao at Vectara. Google didn’t provide additional comment. For the leaderboard’s purposes, the specific hallucination rate numbers are less important than the overall ranking of each model, says Bao.

    But this ranking may not be the best way to compare AI models.

    For one thing, it conflates different types of hallucinations. The Vectara team pointed out that, although the DeepSeek-R1 model hallucinated 14.3 per cent of the time, most of these were “benign”: answers that are factually supported by logical reasoning or world knowledge, but not actually present in the original text the bot was asked to summarise. DeepSeek didn’t provide additional comment.

    Another problem with this kind of ranking is that testing based on text summarisation “says nothing about the rate of incorrect outputs when [LLMs] are used for other tasks”, says Emily Bender at the University of Washington. She says the leaderboard results may not be the best way to judge this technology because LLMs aren’t designed specifically to summarise texts.

    These models work by repeatedly answering the question of “what is a likely next word” to formulate answers to prompts, and so they aren’t processing information in the usual sense of trying to understand what information is available in a body of text, says Bender. But many tech companies still frequently use the term “hallucinations” when describing output errors.

    “‘Hallucination’ as a term is doubly problematic,” says Bender. “On the one hand, it suggests that incorrect outputs are an aberration, perhaps one that can be mitigated, whereas the rest of the time the systems are grounded, reliable and trustworthy. On the other hand, it functions to anthropomorphise the machines – hallucination refers to perceiving something that is not there [and] large language models do not perceive anything.”

    Arvind Narayanan at Princeton University says that the issue goes beyond hallucination. Models also sometimes make other mistakes, such as drawing upon unreliable sources or using outdated information. And simply throwing more training data and computing power at AI hasn’t necessarily helped.

    The upshot is, we may have to live with error-prone AI. Narayanan said in a social media post that it may be best in some cases to only use such models for tasks when fact-checking the AI answer would still be faster than doing the research yourself. But the best move may be to completely avoid relying on AI chatbots to provide factual information, says Bender.

    Topics:



    Original Source Link

    Share. Facebook Twitter Pinterest LinkedIn WhatsApp Email Reddit Telegram
    Previous ArticleI May Never Get Over Thunderbolts*’ 10 Most Jaw-Dropping MCU Movie Moments
    Next Article $25 Off DoorDash Promo Code | May 2025

    RELATED POSTS

    The Data Center Backlash Should Also Be a Climate Reckoning. It Isn’t Yet

    September 27, 2026

    How a horse flu crisis fueled animals rights advocacy in the U.S.

    September 27, 2026

    A Gravitational Battle Within the Earth Is Changing the Length of Days

    September 26, 2026

    Scientists just found a bizarre cosmic object ‘hiding in plain sight’

    September 26, 2026

    Why This Weekend’s Nor’easter Is Like a Hurricane

    September 25, 2026

    Trump’s FDA nominee Heidi Overton affirms MMR vaccine and mifepristone are safe at crucial Senate hearing

    September 25, 2026
    latest posts

    Blink-182’s Mark Hoppus reveals how he sang on unreleased Linkin Park song

    Blink-182’s Mark Hoppus has revealed that he recorded vocals for an unreleased Linkin Park song…

    Coastal erosion is so bad in California that this mostly GOP city may tax itself to restore beach

    September 27, 2026

    Trump Can’t Get Even Go To The Golf Course Without Being Called A Pedo

    September 27, 2026

    Brain fog and memory loss may signal reversible causes, neurologist says

    September 27, 2026

    Anthropic’s Dario Amodei gets the SNL treatment

    September 27, 2026

    The Data Center Backlash Should Also Be a Climate Reckoning. It Isn’t Yet

    September 27, 2026

    Sophie Thatcher: ‘It felt like the most I’d ever…

    September 27, 2026
    Categories
    • Books (1,516)
    • Business (6,420)
    • Events (77)
    • Film (6,354)
    • Lifestyle (4,419)
    • Music (6,484)
    • Politics (6,408)
    • Science (5,772)
    • Technology (6,353)
    • Television (6,045)
    • Uncategorized (9)
    • US News (6,407)
    popular posts

    “This Or That” With Debut Novelist Stephanie Burns

    Debut novelist Stephanie Burns admits she’s been consumed by pop culture for as long as…

    The 15 Best Horror Movies On Peacock, Ranked

    November 23, 2023

    Riverdale Season 7 Episode 14 Review: Chapter One Hundred Thirty-One: Archie the Musical

    July 6, 2023

    Rapper Desiigner Charged with Indecent Exposure on Plane

    April 25, 2023
    Archives
    Browse By Category
    • Books (1,516)
    • Business (6,420)
    • Events (77)
    • Film (6,354)
    • Lifestyle (4,419)
    • Music (6,484)
    • Politics (6,408)
    • Science (5,772)
    • Technology (6,353)
    • Television (6,045)
    • Uncategorized (9)
    • US News (6,407)
    About Us

    We are a creativity led international team with a digital soul. Our work is a custom built by the storytellers and strategists with a flair for exploiting the latest advancements in media and technology.

    Most of all, we stand behind our ideas and believe in creativity as the most powerful force in business.

    What makes us Different

    We care. We collaborate. We do great work. And we do it with a smile, because we’re pretty damn excited to do what we do. If you would like details on what else we can do visit out Contact page.

    Our Picks

    The Data Center Backlash Should Also Be a Climate Reckoning. It Isn’t Yet

    September 27, 2026

    Sophie Thatcher: ‘It felt like the most I’d ever…

    September 27, 2026

    Bill Maher & Roseanne Barr Eviscerate ‘Insane’ Far Left — Then Things Get Heated

    September 27, 2026
    © 2026 New York Examiner News. All rights reserved. All articles, images, product names, logos, and brands are property of their respective owners. All company, product and service names used in this website are for identification purposes only. Use of these names, logos, and brands does not imply endorsement unless specified. By using this site, you agree to the Terms & Conditions and Privacy Policy.

    Type above and press Enter to search. Press Esc to cancel.

    We use cookies on our website to give you the most relevant experience by remembering your preferences and repeat visits. By clicking “Accept All”, you consent to the use of ALL the cookies. However, you may visit "Cookie Settings" to provide a controlled consent.
    Cookie SettingsAccept All
    Manage consent

    Privacy Overview

    This website uses cookies to improve your experience while you navigate through the website. Out of these, the cookies that are categorized as necessary are stored on your browser as they are essential for the working of basic functionalities of the website. We also use third-party cookies that help us analyze and understand how you use this website. These cookies will be stored in your browser only with your consent. You also have the option to opt-out of these cookies. But opting out of some of these cookies may affect your browsing experience.
    Necessary
    Always Enabled
    Necessary cookies are absolutely essential for the website to function properly. These cookies ensure basic functionalities and security features of the website, anonymously.
    CookieDurationDescription
    cookielawinfo-checkbox-analytics11 monthsThis cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Analytics".
    cookielawinfo-checkbox-functional11 monthsThe cookie is set by GDPR cookie consent to record the user consent for the cookies in the category "Functional".
    cookielawinfo-checkbox-necessary11 monthsThis cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".
    cookielawinfo-checkbox-others11 monthsThis cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Other.
    cookielawinfo-checkbox-performance11 monthsThis cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Performance".
    viewed_cookie_policy11 monthsThe cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.
    Functional
    Functional cookies help to perform certain functionalities like sharing the content of the website on social media platforms, collect feedbacks, and other third-party features.
    Performance
    Performance cookies are used to understand and analyze the key performance indexes of the website which helps in delivering a better user experience for the visitors.
    Analytics
    Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics the number of visitors, bounce rate, traffic source, etc.
    Advertisement
    Advertisement cookies are used to provide visitors with relevant ads and marketing campaigns. These cookies track visitors across websites and collect information to provide customized ads.
    Others
    Other uncategorized cookies are those that are being analyzed and have not been classified into a category as yet.
    SAVE & ACCEPT