Close Menu
New York Examiner News

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Lainey Wilson Set to Host 2026 CMA Awards

    September 5, 2026

    U.S. sanctions Turkish bank, accusing it of helping Iran transfer oil revenue from China to Turkey

    September 5, 2026

    Arizona Rancher Tells Trump That Dem Gov Katie Hobbs Wants to Evict Him From Family Land to Build Foreign Company’s Solar Energy Plant During Oval Office Meeting * The Gateway Pundit * by Jordan Conradson

    September 5, 2026
    Facebook X (Twitter) Instagram
    New York Examiner News
    • Home
    • US News
    • Politics
    • Business
    • Science
    • Technology
    • Lifestyle
    • Music
    • Television
    • Film
    • Books
    • Contact
      • About
      • Amazon Disclaimer
      • DMCA / Copyrights Disclaimer
      • Terms and Conditions
      • Privacy Policy
    New York Examiner News
    Home»Business»OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch
    Business

    OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

    By AdminSeptember 5, 2026
    Facebook Twitter Pinterest LinkedIn WhatsApp Email Reddit Telegram
    OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch



    OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3. In some cases, the numbers on the updated versions showed Astra performing better, while numbers for models from OpenAI’s arch rival Anthropic got worse.

    The changes occurred amid an unusual rollout of the blog post. OpenAI originally planned for the post to go live at 2 p.m. ET, but it took almost another two hours before it was widely viewable online.

    When OpenAI’s X account tweeted out the blog post at 3:32 p.m., the link was not loading properly, returning an error message. At 3:50 p.m., OpenAI CEO Sam Altman posted the link, writing, “We hit a little snag getting the blog post deployed, but it is really great.” Multiple commenters were still unable to see it, and were getting the same error, as did Fortune. When we checked back about an hour later, it was visible and loading properly.

    It turns out OpenaAI actually published the blog shortly after 2pm but retracted it for reason the company said it could not disclose, but which it said were unrelated to the benchmark performance figures. (OpenAI first told us it was a bug in the content management system, and then an internet outage.) Upon republishing the blog, it had different evaluation metrics that seemed to favor Astra—and some figures have continued to change even since then.

    The revelation of the changes comes amid intense competition in the AI industry, as companies release updates to their large language models at a frenetic pace, each seeking to pull ahead of the other. The focus on metrics also highlights the challenges of measuring the performance of large language models using standardized benchmark tests and concerns that the specs are prone to manipulation and gamesmanship.

    “We care deeply about getting evaluations right,” an OpenAI spokesperson told Fortune. “Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.”

    Discrepancies between the first and final published blogs—and the numbers are still changing

    Among the most notable changes was Astra’s reported hallucination rate. In the first internet archive snapshot of the blog post from 2:23 p.m. ET, it was 4.2%. It remained that number for several more snapshots, the last being a fifth at 3:11 p.m. ET—about 10 minutes before OpenAI tweeted out the final version.

    But the hallucination rate, along with four other metrics, changed in the sixth archival snapshot of the page taken at 5:20 p.m.—after everyone could likely finally see the blog. It was halved down to 2% for Astra. The scores for Astra’s predecessor, GPT-5.6 Sol, also went down from 12.2% to 9.4%. OpenAI has continued to change this metric; as of this writing, the hallucination rates are back up to their original 4.2% and 12.2%.

    OpenAI also seems to have given GPT-5.6 Sol a big boost on its internal version of the ExploitBench cybersecurity evaluation, going from 5.5% in the first version to 11.5% in the later versions. OpenAI said it is currently investigating reverting that number back to 5.5% because it says the 11.5% result reflects a reasoning level that is not commercially available for Sol.

    Astra is especially good at mathematics, OpenAI says, a quality the company highlights in the opening paragraph of the announcement page. While that metric did not change in the snapshots for Astra—it stays at 97.6% for the FrontierMath Tier 4 (v2) eval—OpenAI did briefly alter the scores for GPT-5.6 Sol and Anthropic’s latest model, Fable 5.1.

    The result of these changes made Astra briefly appear significantly better at math than those two models. In the first snapshot (2:23 p.m. on Sept. 3), Anthropic’s Fable 5.1 model’s score is 87.8%. By 5:17 p.m., it’s dropped nearly 10 percentage points to 78%. Today, it’s back up to 83%. Similarly, GPT-5.6 Sol’s scores go from 83%, down to 80.5%, and back up to 83% today.

    The changes in metrics began even before OpenAI first published its blog at 2 p.m. An embargoed pre-publication draft the company provided to Fortune and other media organizations listed Astra’s score on the ARC-AGI-3 evaluation as 98.6%. It’s now 99.99% in the live blog.

    “We always verify evals before publication so adjustments between draft and final version are normal,” a company spokesperson said at the time. OpenAI also noted that the creator of the benchmark, the Arc Prize Foundation, found that Astra performed at 99.9% in its independent assessment, provided the model was given a particularly powerful harness (a set of tools the model can use to complete tasks). It performed at 63%—still significantly better than any other AI model currently in public release—when given the benchmark’s standard harness. OpenAI said “things like harness, reasoning level and other factors inform evals.”

    “Benchmaxxing”—or improving accuracy?

    Different research teams at OpenAI oversee different metrics, and are responsible for calculating and reporting them to a central team to publish. OpenAI is open about the fact that the numbers are achieved under the best possible conditions and may be slightly different from the models available in the production ChatGPT product that most users can access. “Evaluation scores are the maximum at any effort,” reads a disclaimer on the blog. The company includes further caveats on each metric in footnotes.

    Accuracy is elusive, as multiple numbers can be considered accurate based on the conditions in which the tests occurred. But some AI experts wonder if there’s also “benchmaxxing” involved. This is a known practice in the AI industry—not just at OpenAI—to maximizing scores by re-running evaluations with different conditions.

    “This can be done in a very tight timeframe, and it’s better for their marketing,” said Anka Reuel and Mike Hardy, researchers at the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab. They also pointed out that the GPT-6 Astra system card, which should contain more technical information on how the evaluations were performed, does not always properly explain them. For the internal hallucination benchmark, for example, the system card provides “barely any details about the evaluation,” they said. “It doesn’t even include the number of test items.”

    This re-running of the numbers could be why Astra’s coding capabilities also got a marginal boost in the later versions of the blog post, up from 57.7% to 57.9%. Though it’s a negligible difference, OpenAI seemed to care enough about it to swap in the new and improved number.

    Not all changes OpenAI made portrayed Astra more favorably. For example, two Anthropic model scores improve in the different versions of the healthcare-focused eval HealthBench Professional. Claude Fable 5.1 goes from 56.6% to 58.1%, and Opus 5 goes from 54.5% to 56.4%. The scores for models made by other AI companies are usually taken from published leaderboards and do not involve OpenAI itself running assessments on rivals’ models.

    Evaluation score debates haunt the AI industry

    The question of benchmark accuracy has come up multiple times in the past. In 2025, Meta denied reports that it artificially boosted scores for its Llama 4 model by publishing results from an internal version of the model rather than the one it was making publicly-available. Yann LeCun, the former chief AI scientist at Meta, later admitted that the company had “fudged” the benchmark results.

    Evaluation metrics also change frequently, as new ones get created. For example, ExploitGym, a cybersecurity benchmark that was at the center of the July incident in which OpenAI’s models went rogue and attacked the company Hugging Face, was created in 2026.

    Vincent Sunn Chen, an AI engineer who leads benchmark and evaluation research at Snorkel AI, said that it’s not unusual for benchmark scores to shift in the final hours before a model launches. “It’s usually a function of final launch logistics,” he said in an email. “A benchmark score reflects a specific measurement setup: the model checkpoint, configuration (including how much time and compute the model is allowed), harness, eval/grading configuration (e.g., non-determinism in the judge). All of those are typically still shifting in the final days before a launch, so I’m not surprised that there were some updates.”

    He said he would like to see industry norms developed that companies should report what has changed about the assessment when a company revises benchmark performance numbers so that researchers can interpret the results more clearly.

    Benchmark results matter for several reasons. They are the way AI companies measure progress—but also a way to keep score in the race against competing AI companies. Topping the leaderboards for these evaluations can help AI companies win customers, and in some cases help them hire engineers and researchers.

    But as this example illustrates, interpreting the benchmark scores can be technically complex, presenting a challenge for companies that want to show off the results to the public in a digestible format. These complexities, as well as confusion over changing metrics and accusations that companies have not been intellectually honest in how they’ve presented the results, could make it difficult for customers and investors to figure out exactly which models are best for which tasks. The confusion could muddy the narrative of having the best models in the market that OpenAI would no doubt like to present ahead of a possible 2027 IPO.



    Original Source Link

    Share. Facebook Twitter Pinterest LinkedIn WhatsApp Email Reddit Telegram
    Previous ArticleTrump Melted Down After A Reporter Challenged Him On The Iran War
    Next Article Nominate your favourite UK grassroots music venue to win a bundle of Fender gear

    RELATED POSTS

    U.S. sanctions Turkish bank, accusing it of helping Iran transfer oil revenue from China to Turkey

    September 5, 2026

    Josh Kushner: Thrive would have stayed away from World Cup deal if it had known what was coming

    September 4, 2026

    GAO finds Secret Service left drone threats unaddressed before Trump assassination attempt

    September 4, 2026

    AI wants electricity now. The electric grid needs years to catch up

    September 3, 2026

    China demands answers after Chinese man dies in ICE custody, the fifth to die in U.S. custody

    September 3, 2026

    Despite being a multimillionaire, Suze Orman says eating out is one of the biggest wastes of money

    September 2, 2026
    latest posts

    Lainey Wilson Set to Host 2026 CMA Awards

    Lainey Wilson is set to become the first woman to solo-host the CMA Awards twice.…

    U.S. sanctions Turkish bank, accusing it of helping Iran transfer oil revenue from China to Turkey

    September 5, 2026

    Arizona Rancher Tells Trump That Dem Gov Katie Hobbs Wants to Evict Him From Family Land to Build Foreign Company’s Solar Energy Plant During Oval Office Meeting * The Gateway Pundit * by Jordan Conradson

    September 5, 2026

    Texas opens season against Texas State as Arch Manning faces a make-or-break college football year

    September 5, 2026

    Oura is going public, but these smart ring companies are coming for its crown

    September 5, 2026

    Scientists Put Caterpillars in an Ultraquiet Chamber to Learn How They Hear Without Ears

    September 5, 2026

    18 Years Later, A Beloved PlayStation 2 JRPG Is Coming To PS5

    September 5, 2026
    Categories
    • Books (1,472)
    • Business (6,375)
    • Events (72)
    • Film (6,310)
    • Lifestyle (4,383)
    • Music (6,438)
    • Politics (6,361)
    • Science (5,727)
    • Technology (6,308)
    • Television (5,998)
    • Uncategorized (9)
    • US News (6,363)
    popular posts

    ‘Diversity’ Doesn’t Include Disabled Veterans Like Me

    Sklifosovsky Insitute, CC BY 4.0, via Wikimedia Commons By Matthew Winans for RealClearPolitics At college…

    ‘The Challenge’ Season 40 Exit Interview: Aneesa Ferreira

    September 5, 2024

    Chargers legend Shawne Merriman says Bill Belichick would be his ‘last’ pick for coaching job in Los Angeles

    January 4, 2024

    Live Nation Entertainment Faces Wage Theft Lawsuit in California

    September 9, 2023
    Archives
    Browse By Category
    • Books (1,472)
    • Business (6,375)
    • Events (72)
    • Film (6,310)
    • Lifestyle (4,383)
    • Music (6,438)
    • Politics (6,361)
    • Science (5,727)
    • Technology (6,308)
    • Television (5,998)
    • Uncategorized (9)
    • US News (6,363)
    About Us

    We are a creativity led international team with a digital soul. Our work is a custom built by the storytellers and strategists with a flair for exploiting the latest advancements in media and technology.

    Most of all, we stand behind our ideas and believe in creativity as the most powerful force in business.

    What makes us Different

    We care. We collaborate. We do great work. And we do it with a smile, because we’re pretty damn excited to do what we do. If you would like details on what else we can do visit out Contact page.

    Our Picks

    Scientists Put Caterpillars in an Ultraquiet Chamber to Learn How They Hear Without Ears

    September 5, 2026

    18 Years Later, A Beloved PlayStation 2 JRPG Is Coming To PS5

    September 5, 2026

    Their Affair, Marriage, Kids, and More

    September 5, 2026
    © 2026 New York Examiner News. All rights reserved. All articles, images, product names, logos, and brands are property of their respective owners. All company, product and service names used in this website are for identification purposes only. Use of these names, logos, and brands does not imply endorsement unless specified. By using this site, you agree to the Terms & Conditions and Privacy Policy.

    Type above and press Enter to search. Press Esc to cancel.

    We use cookies on our website to give you the most relevant experience by remembering your preferences and repeat visits. By clicking “Accept All”, you consent to the use of ALL the cookies. However, you may visit "Cookie Settings" to provide a controlled consent.
    Cookie SettingsAccept All
    Manage consent

    Privacy Overview

    This website uses cookies to improve your experience while you navigate through the website. Out of these, the cookies that are categorized as necessary are stored on your browser as they are essential for the working of basic functionalities of the website. We also use third-party cookies that help us analyze and understand how you use this website. These cookies will be stored in your browser only with your consent. You also have the option to opt-out of these cookies. But opting out of some of these cookies may affect your browsing experience.
    Necessary
    Always Enabled
    Necessary cookies are absolutely essential for the website to function properly. These cookies ensure basic functionalities and security features of the website, anonymously.
    CookieDurationDescription
    cookielawinfo-checkbox-analytics11 monthsThis cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Analytics".
    cookielawinfo-checkbox-functional11 monthsThe cookie is set by GDPR cookie consent to record the user consent for the cookies in the category "Functional".
    cookielawinfo-checkbox-necessary11 monthsThis cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".
    cookielawinfo-checkbox-others11 monthsThis cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Other.
    cookielawinfo-checkbox-performance11 monthsThis cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Performance".
    viewed_cookie_policy11 monthsThe cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.
    Functional
    Functional cookies help to perform certain functionalities like sharing the content of the website on social media platforms, collect feedbacks, and other third-party features.
    Performance
    Performance cookies are used to understand and analyze the key performance indexes of the website which helps in delivering a better user experience for the visitors.
    Analytics
    Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics the number of visitors, bounce rate, traffic source, etc.
    Advertisement
    Advertisement cookies are used to provide visitors with relevant ads and marketing campaigns. These cookies track visitors across websites and collect information to provide customized ads.
    Others
    Other uncategorized cookies are those that are being analyzed and have not been classified into a category as yet.
    SAVE & ACCEPT