AI-assisted safety marketing campaign centered on the Bitcoin ecosystem, Bitcoin Crimson Workforce, mentioned it generated 6,700 findings throughout 425 initiatives in its first 55 hours. The marketing campaign labeled 1,029 of them excessive or crucial.
The Aug. 6 replace measures how a lot materials entered a safety triage pipeline, and its impact on software program safety stays unreported.
The retrieved thread omitted audit-ready definitions and denominators for the severity counts, in addition to case-level outcomes, an mixture false-positive fee, and a repair fee.
These lacking fields forestall a calculation of what number of alerts turned confirmed vulnerabilities, what number of maintainers rejected or downgraded, and what number of led to patches.
The primary 55 hours nonetheless reveal a consequential functionality, noting how AI methods can fill an ecosystem-scale assessment pipeline rapidly. Knowledgeable prompting, replica, disclosure, and maintainer response remained mandatory at each later stage.
What the marketing campaign numbers measure
The marketing campaign printed two snapshots as its roster and workload expanded:
Elapsed timeProjectsTotal findingsReported severityParticipants27.5 hours3904,96285 crucial; 635 high1655 hours4256,7001,029 excessive or critical24 reported, together with three bots
The 27.5-hour replace coated 390 initiatives and 4,962 findings. By the 55-hour mark, the challenge rely had risen by 35 and the discovering rely by 1,738. The later thread put high-or-critical findings at 15.4% of the whole and clarified that three of the 24 reported individuals had been bots.
The sooner submit separated crucial and excessive findings, whereas the later one mixed them, with each units of figures reflecting marketing campaign assessments. Maintainer-confirmed exploitability and remediation outcomes require separate proof.
Rob Hamilton described Kimi K3 as dealing with the heavy evaluation, with GPT Sol, Fable/Opus, and GLM 5.2 supporting the documentation. He mentioned OpenAI’s Cyber Harness coated chosen elements he thought of load-bearing.
A day later, Hamilton wrote that subject-matter consultants might change an evaluation with one or two sentences of context or a small block of code. In examples he described, that enter pushed middling issues into excessive or crucial territory. He additionally recognized operations, disclosure handoff, and triage as bottlenecks.
In Hamilton’s account, fashions searched broadly whereas specialists formed prompts, interpreted output, tried replica, and determined which stories had been prepared for disclosure. That division of labor makes the marketing campaign a human-AI assessment system.
The developer often known as Calle mentioned most important stories had been rapidly verified by challenge house owners. The submit equipped no denominator, verified-report rely, rejection rely, or patch standing, leaving the breadth and consequence of that verification unresolved.
Outreach and outcomes outline the safety worth
Within the 55-hour replace, Bitcoin Crimson Workforce reported that 19.5% of scanned initiatives had a SECURITY.md file and 13.1% had an electronic mail there. The retrieved thread omitted the challenge corpus, denominator interpretation, and measurement methodology, so the odds solely describe the marketing campaign’s scan.
On Aug. 3, Hamilton mentioned the trouble had spent over $10,000 scanning over 100 repositories and had instantly disclosed crucial findings when a proof of idea demonstrated exploitability. On Aug. 4, he reported about $20,000 in spending, greater than a dozen disclosures and 150 repositories scanned.
Scanning continued to develop, whereas the marketing campaign described outreach, handoff and triage as lively operational constraints. The printed snapshots provide no comparable disclosure denominator at 55 hours, so they can’t set up the relative pace of scanning and determination.
Hamilton later recognized the separate Coldcard incident as a catalyst for the broader marketing campaign. The marketing campaign report attributes no discovery of the Coldcard flaw to this dash.
A helpful public accounting would separate findings that had been reproduced, acknowledged, downgraded, rejected, and glued, with definitions and denominators for every fee. That breakdown would present how a lot of the marketing campaign’s quantity turned actionable safety work.
A public critic, JW Weatherman, argued that the marketing campaign couldn’t triage its output. His submit recognized no campaign-linked concern, patch, or advisory, so it provides criticism and not using a measurable failure fee. The marketing campaign’s lacking disposition knowledge leaves the underlying query open.
For now, 6,700 represents campaign-labeled findings and triage candidates. The dash demonstrated the pace of machine-assisted assessment. Its lasting safety worth is dependent upon the share that consultants can validate, disclose, and convert into fixes.








