AI Sprint Flags 6,700 Issues Across 425 Bitcoin Projects

AI-assisted Bitcoin Red Team reported 6,700 findings across 425 projects in 55 hours, with 1,029 labeled high or critical; verification and fix rates were not published.

Bitcoin Red Team, an AI-assisted security campaign, reported 6,700 findings across 425 Bitcoin-related projects during its first 55 hours of scanning. The campaign labeled 1,029 of those findings as high or critical and did not publish verification or fix rates.

The campaign published two progress snapshots. After 27.5 hours it reported scanning 390 projects and producing 4,962 findings, including 85 critical and 635 high issues. At 55 hours the roster had grown to 425 projects and 6,700 total findings, with 1,029 listed as high or critical. The later update named 24 participants or accounts, three of which were identified as bots.

Organizers described a mix of automated model runs and human review. Rob Hamilton wrote that a model called Kimi K3 performed the bulk of analysis, while GPT Sol, Fable/Opus and GLM 5.2 supported documentation. Hamilton also noted that OpenAI’s Cyber Harness was used on selected components. He wrote, “subject-matter experts could change an assessment with one or two sentences of context or a small block of code.” Specialists shaped prompts, interpreted outputs, attempted reproduction and decided which reports were ready to disclose.

A developer using the name Calle reported that most critical reports were quickly verified by project owners. Hamilton wrote that the team immediately disclosed issues when a proof of concept demonstrated exploitability. He also identified outreach, disclosure handoff and triage as operational bottlenecks. Hamilton referenced a separate Coldcard incident as a catalyst for the wider sprint but did not attribute discovery of that flaw to the sprint itself.

The campaign published a small set of operational metrics. In the 55-hour update Bitcoin Red Team reported that 19.5% of scanned projects included a SECURITY.md file and 13.1% listed a contact email. The post did not provide the underlying project list or the scan method used to identify those files. Hamilton reported spending more than $10,000 by Aug. 3 to scan over 100 repositories and about $20,000 by Aug. 4, when the team had disclosed more than a dozen issues and scanned roughly 150 repositories.

Key disposition details were not published. The campaign’s public updates did not include definitions for severity levels, a verified-report count, an aggregate false-positive rate, a rejection count, patch status or case-level outcomes. Those omissions prevent calculation of how many flagged items became confirmed vulnerabilities, how many maintainers disagreed or downgraded findings, and how many led to patches.

A public critic, JW Weatherman, argued the campaign could not triage its output; his post did not identify campaign-linked advisories or patches. The Bitcoin Red Team’s updates described the 6,700 items as labeled findings and triage candidates rather than confirmed or remediated vulnerabilities.

As of the last update, the campaign reported the scale of scanning activity and the mix of automated and human review but did not publish verification or fix rates that would show case-level outcomes.

Articles by this author