Home / Publications / Blog / ForIntel Research / Apple Just Rationed Its Own Bug Bounty Because It Couldn't Tell Real Bugs From AI Hallucinations

Apple Just Rationed Its Own Bug Bounty Because It Couldn't Tell Real Bugs From AI Hallucinations

In 2026 Apple capped how many security reports a researcher can open at once after AI-generated, hallucinated vulnerability claims flooded its review pipeline — and a real six-figure macOS flaw nearly went unreported as a result.

By Foragentis teamPublished 2026-08-056 min read

In June 2026, Apple made a change to its own Security Bounty program that says more about the state of AI-generated content than any retracted consulting report has: it started rationing submissions. Researchers can now only have a capped number of vulnerability reports open with Apple at once, and once they hit that cap, they face a mandatory 30-day cooldown before they can submit more (a higher quota can be requested, but it isn't automatic). Apple's own explanation, given to reporters: "With the growing volume of AI-generated security submissions across the industry, we recently adjusted the number of new reports a researcher can open at once."

Why would Apple cap its own bug bounty?

Because the review pipeline was drowning. The problem isn't that AI is bad at finding bugs. It's that large language models now let researchers with limited technical skill produce large volumes of vulnerability reports that read as plausible, technically fluent security findings — and a meaningful share of them describe exploits that don't actually work, or don't exist at all. Every one of those still needs a human reviewer to work out whether it's a real flaw or a hallucination dressed up in the right vocabulary. The volume got bad enough that Apple's answer wasn't better filtering. It was a hard cap on how much could come in.

A hallucinated vulnerability report is a security submission that names real components, uses correct terminology, and follows the expected format of a genuine finding, but describes an exploit that cannot be reproduced because the underlying flaw was never there. It is fluent enough to survive a first read and false enough to waste the reviewer who takes it seriously.

Apple already tried the smarter fix. It wasn't enough.

This sits on top of, and is a separate move from, Apple's October 2025 Security Bounty overhaul, which introduced "Target Flags": a capture-the-flag-style requirement where a researcher has to prove a claimed exploit actually reaches a protected part of the system before Apple pays out, alongside a top payout raised past $5 million, the largest in the industry. That change made genuine, provable findings faster to verify and better paid. It did not stop the flood of fabricated ones. Eight months later, Apple needed a second, blunter control just to keep the pipeline moving.

The cost: a real six-figure flaw that nearly went unreported

The story's clearest illustration of the cost, reported by the Financial Times on August 2, 2026: Bynario, a seven-person Italian security startup led by CEO Alfredo Pesoli, used ChatGPT to find a genuine macOS privilege-escalation exploit chain, the kind of flaw that can hand an attacker full control of a machine. Bynario tried to report it and couldn't, because the company had already hit Apple's new submission cap. Pesoli estimated the flaw's black-market value at $100,000 to $200,000. Apple reportedly reached out to Bynario directly once the Financial Times reporting surfaced the gap. Sophos threat researcher Rafe Pilling summed up what's changed in bug-bounty triage industry-wide: it has shifted from finding vulnerabilities to "validating them at machine speed."

Read that sequence again: a real, six-figure vulnerability nearly went unreported, not because nobody found it, but because the review pipeline was too clogged with fabricated ones to accept it in time. That's not an AI-capability failure. It's a verification failure, and it happened to one of the best-resourced security organizations on earth.

The same missing layer, in a different industry

The bug-bounty flood is the security-research version of a pattern we've been tracking elsewhere. When all four of the Big Four consulting firms published AI-hallucinated research reports, the shared root cause wasn't the AI tool — it was the absence of a step that checks every claim against a real, verifiable source before publication. Apple's flooded pipeline is the same gap wearing a different uniform: fluent, format-correct output that no one has verified, arriving faster than any human can validate it. A platform can be one of the most sophisticated on earth and still get overwhelmed the moment plausibility becomes cheap to manufacture.

Where the durable fix actually sits

At Foragentis, this is the exact problem our intelligence architecture is built to prevent, one step earlier than Apple's fix. Target Flags force a researcher to prove reachability before Apple pays; every claim in a Foragentis report has to resolve to a verifiable source row before it's allowed to publish at all. Acceptance criteria are pre-registered before drafting starts, not applied after something looks off. Every report includes a known-limits section naming what couldn't be verified, so a claim never gets to masquerade as fact just because it's fluent. Apple had to solve this with volume caps and a cooldown timer — a rationing fix. The more durable fix is verification at the point of intake, before a claim is ever competing for a human's attention in the first place.

This is exactly the failure mode ForIntel was built to catch: foragentis.com.

FAQ

Why did Apple cap its bug bounty submissions in 2026?

Apple limited how many vulnerability reports a researcher can have open at once because AI-generated security submissions had flooded its review pipeline. Large language models let researchers with limited technical skill produce high volumes of plausible-looking vulnerability reports, a meaningful share of which describe exploits that don't work or don't exist. Every one still needs a human reviewer to sort real flaws from hallucinations. In Apple's own words, given to reporters: "With the growing volume of AI-generated security submissions across the industry, we recently adjusted the number of new reports a researcher can open at once."

How does Apple's Security Bounty submission cap work?

Under the 2026 change, a researcher can only have a capped number of vulnerability reports open with Apple at any one time. Once they hit that cap, they face a mandatory 30-day cooldown before submitting more. A higher quota can be requested, but it is not granted automatically.

What happened with Bynario's macOS exploit?

As the Financial Times reported on August 2, 2026, Bynario — a seven-person Italian security startup led by CEO Alfredo Pesoli — used ChatGPT to find a genuine macOS privilege-escalation exploit chain, the kind of flaw that can hand an attacker full control of a machine. Bynario tried to report it to Apple and couldn't, because the firm had already hit Apple's new submission cap. Pesoli estimated the flaw's black-market value at $100,000 to $200,000. Apple reportedly reached out to Bynario directly once the FT reporting surfaced the gap.

What are Apple's "Target Flags"?

Target Flags are a capture-the-flag-style requirement Apple introduced in its October 2025 Security Bounty overhaul: a researcher has to prove a claimed exploit actually reaches a protected part of the system before Apple pays out. Apple paired it with a top payout raised past $5 million, the largest in the industry. Target Flags made genuine, provable findings faster to verify and better paid, but they did not stop the flood of fabricated submissions — which is why Apple added the blunter volume cap eight months later.

How does Foragentis prevent AI-hallucinated claims from reaching a reader?

Every claim in a Foragentis report has to resolve to a verifiable source row in a dedicated data layer before it is allowed to publish. Acceptance criteria are pre-registered before drafting begins, not applied after something looks off, and every report includes a published "known limits" section naming what could not be verified. The verification happens at the point of intake — before a claim ever competes for a human reviewer's attention — rather than as a volume cap applied after the pipeline is already clogged.


ForIntel is a research-first business-intelligence product from Foragentis. Every quantitative claim in a ForIntel brief traces to source data, is cross-corroborated across independent measures, and passes a counter-signal verification step before it is published. To see a sample brief or commission a read, reach the ForIntel desk at forintel@foragentis.com.