Anthropic’s Open-Source Tool Lets Claude Hunt and Fix Code Flaws on Autopilot
Anthropic dropped a new open-source toolkit this week that could change how security teams handle vulnerable code. The repository, posted on GitHub, gives developers a ready-made system for using Claude to find and patch security holes automatically.
The release comes as research shows AI coding assistants are flooding repositories with functional but insecure code. Veracode’s latest report found only 55% of AI-generated code passes security checks, even though syntax accuracy tops 95%. The gap between what works and what’s safe remains wide.
But spotting bugs isn’t the hard part anymore. Stronger models can do that at scale. The real bottlenecks are verifying, prioritizing, and fixing issues. Anthropic’s new harness targets those exact steps.
The system includes tools for threat modeling, scanning, triage, and patching. Users can run a single command to start an interactive session or set the harness loose to scan continuously. Everything runs inside Claude Code, Anthropic’s agentic coding environment.
Threat modeling comes first. The tool builds a draft from your code, known vulnerabilities, and git history. Then it walks you through four questions to refine the model. The output becomes a living document that later stages use to guide discovery and patching.
Discovery follows, with Claude scanning based on that threat model. It flags potential issues and collects evidence. Triage ranks findings by severity and likelihood, avoiding the common trap where every alert causes panic—or none does.
Patching closes the loop. The harness generates fixes, tests them, and proposes changes. It supports Docker and AddressSanitizer for memory bugs in C and C++. Git history helps it spot recurring patterns.
The foundation comes from Project Glasswing, Anthropic’s initiative to secure critical software. The company worked with security teams at several organizations to develop the approach. Now they’re sharing it as an open-source reference.
Security practitioners on X and developer forums reacted quickly. Many praised the structured workflow and connection to established practices like threat modeling. Some cautioned that autonomous remediation still needs human oversight. The harness amplifies judgment, it doesn’t replace it.
Recent studies underscore the need. Opsera’s 2026 benchmark found AI cut time-to-pull-request by 58%, but those PRs waited 4.6 times longer in review and introduced 15 to 18% more vulnerabilities. Speed without control creates debt.
Anthropic’s approach focuses on the full cycle. The accompanying blog post walks through each stage and invites users to run the quickstart command. Teams can fork the repo, customize the skills, and connect it to their CI pipelines.
Enterprises should start small. Test the quickstart on a non-production codebase. Examine the generated threat model. Only scale to autonomous operation once the process is solid. The harness makes security repeatable. It doesn’t make it automatic.
The timing feels deliberate. As more companies adopt agentic coding, the attack surface expands. Defenders need equally capable systems. This harness gives them a strong opening move.
Source: Webpronews
dArt Studio installs AI for local businesses in Broward & Palm Beach County, FL. We reply within 1 business hour.