Tag
ai-safety
12 articles

Strix: The Open-Source AI Pentester That Writes Exploits
Strix is an open-source (Apache-2.0) AI penetration-testing tool with ~39,000 GitHub stars. Its autonomous agents dynamically run your app, exploit OWASP Top 10 vulnerabilities, and validate each finding with a working proof-of-concept, cutting the false positives of static scanners. It installs via a single curl command, needs Docker plus an LLM API key, is model-agnostic through LiteLLM, and drops into CI/CD with a non-interactive mode that fails builds on findings.
By Marcus Rivera · 5 min · Jul 24, 2026

Project Perception: Microsoft's Cheaper Rival to Claude Mythos
Microsoft is reportedly developing Project Perception, a multi-model AI security platform that routes vulnerability-scanning tasks across models from Microsoft, OpenAI, and Anthropic to reserve expensive frontier calls for high-value steps. Its pitch is matching Anthropic's Claude Mythos on capability while costing far less. Microsoft has not officially confirmed details, so the news should be treated as a credible report pending benchmarks.
By Sarah Chen · 5 min · Jul 21, 2026

TAKE IT DOWN Act: The Deepfake Law Now Binding Every Platform
The TAKE IT DOWN Act's Section 3 set a May 19, 2026 deadline for covered U.S. platforms to offer a removal process for non-consensual intimate images, including AI deepfakes, and to take them down within 48 hours. The FTC enforces it with civil penalties up to $53,088 per violation and has warned 15 major platforms.
By Aisha Patel · 6 min · Jul 16, 2026

DPO: How Direct Preference Optimization Replaced RLHF
Direct Preference Optimization (DPO), introduced in a 2023 NeurIPS paper by Rafailov et al., aligns language models directly on preference pairs without training a separate reward model or running reinforcement learning. It replaces RLHF's fragile four-model PPO pipeline with a single supervised loss governed mainly by one parameter, beta, and works best stacked after SFT on subjective tasks — not on problems with a single correct answer.
By Aisha Patel · 9 min · Jul 13, 2026

AI Hallucinations in Court: 1,725 Cases and a $110K Wake-Up Call
AI hallucinations in court filings have grown from the 2023 Mata v. Avianca case (a $5,000 sanction for six fabricated ChatGPT citations) into a documented worldwide phenomenon. Damien Charlotin's database catalogs 1,725 cases as of July 5, 2026, led by the US (1,187), Canada (190), and Australia (96). Self-represented litigants account for 1,016 cases, lawyers 667. In December 2025, an Oregon federal judge imposed a record $110,000 penalty in Couvrette v. Wisnovsky for 15 fake cases and 8 fabricated quotations. At least 25 federal courts now require AI-use certifications.
By Aisha Patel · 5 min · Jul 7, 2026

New York Kids Chatbot Safety Act: Inside the S9051B Ban
New York's Kids Chatbot Safety Act (S9051B) passed both chambers unanimously in June 2026, banning AI companion chatbots for minors and prohibiting sycophancy and claims of being human. Enforced by the attorney general with fines up to $25,000 per violation, it takes effect January 1, 2027 pending the governor's signature, part of a national wave including California SB 243 and the federal GUARD Act.
By Aisha Patel · 4 min · Jul 6, 2026

Agentjacking: Fake Sentry Errors Hijack Your AI Coding Agent
Agentjacking injects fake Sentry errors that AI coding agents read over MCP as trusted guidance, then execute - hitting an 85% success rate across 2,388 exposed orgs.
By Aisha Patel · 8 min · Jun 29, 2026

EU AI Act: Why Brussels Just Delayed Its Toughest Rules
The EU AI Act's high-risk obligations have been postponed via the Digital Omnibus on AI: stand-alone Annex III systems now apply from 2 December 2027 and embedded Annex I systems from 2 August 2028 (fixed dates, not a conditional trigger). A provisional political deal was struck 6 May 2026 and confirmed by the Council 13 May. A new Article 5 ban on nudifiers/CSAM is added (transition to 2 Dec 2026), and AI literacy duties are softened. Crucially, Article 50 transparency obligations still apply from 2 August 2026. The piece weighs whether the delay is a quiet retreat or responsible governance.
By Aisha Patel · 6 min · Jun 24, 2026

Model Collapse: Why AI Trained on AI Slowly Falls Apart
Model collapse is the progressive degradation of generative models trained recursively on synthetic data, documented in Nature (Shumailov et al., 2024). Errors compound and rare data vanishes, but research (Gerstgrasser et al., 2024) shows accumulating real data alongside synthetic data, tracking ratios, and verifying generations prevents it.
By Aisha Patel · 8 min · Jun 19, 2026

The Great American AI Act: A 3-Year Freeze on State AI Laws
The Great American Artificial Intelligence Act, a 269-page bipartisan discussion draft unveiled June 4, 2026, would impose federal safety mandates on large frontier AI developers—public risk frameworks, semi-annual independent audits, incident reporting, and up to $1 million-a-day penalties—while preempting new state laws regulating AI model development for three years. AI-safety groups call the preemption a 'generational mistake'; sponsors argue a single federal standard beats a 50-state patchwork.
By Aisha Patel · 6 min · Jun 11, 2026

AI Companion Chatbots: The 2026 Lawsuit Reckoning
A 2026 survey of the legal and regulatory reckoning facing AI companion chatbots. Florida sued OpenAI and Sam Altman on June 1, 2026; Character.AI settled teen-suicide suits and faces a Pennsylvania action; the FTC opened a companion-bot inquiry; and the EU AI Act becomes fully applicable on August 2, 2026, but leaves emotion-recognition gaps. The piece outlines what real safeguards would require.
By Aisha Patel · 6 min · Jun 2, 2026

Claude Mythos: The AI Anthropic Built Then Refused to Release
Anthropic trained Claude Mythos, its most capable AI, but refused to release it due to security findings.
By Aisha Patel · 6 min · Apr 18, 2026