September 18, 2026

D.A.D. today covers 8 stories — about a 8-minute read. What's New, What's Innovative, What's Controversial, What's in the Lab, and What's in Academe.

The Daily AI Digest is a daily AI briefing automated by Alexander Panetta — a veteran political journalist tracking the field during a Master's in AI Management at Georgetown University.

D.A.D. Joke of the Day: My company adopted an AI policy this quarter. It has three pages of rules, and the AI wrote all of them.

What's New

AI developments from the last 24 hours

Three Researchers and $3,000 in Tokens Got Into OpenAI's Private Code

On July 25, a three-person team at the security firm Hacktron took over the ChatGPT account of an OpenAI employee and used the coding agent attached to it to open a pull request inside OpenAI's internal code repository. They did it to prove they could, avoided reading anything sensitive, reported it immediately, and were paid a $6,500 bounty. OpenAI shipped a fix in about 14 hours. The Wall Street Journal's Robert McMillan reported the episode Thursday night; the researchers had published their own account on September 13.

The path in is worth following, because none of it involved OpenAI's models. A flaw in libheif — an image-decoding library — let them run code on OpenAI's community help forum by uploading a doctored photo. A separate misconfiguration in OpenAI's single sign-on then turned that foothold into control of forum users' ChatGPT and Codex accounts, and one of those accounts had Codex wired into OpenAI's GitHub. Discovery to repository access took under 72 hours. The underlying bug had been fixed upstream a year earlier, but the fix was never labeled a security fix and never got a CVE number, so Debian never backported it and everything built on it stayed exposed.

What Claude did was write the exploit — and the researchers documented, unusually precisely, when it became able to. Opus 4.8 failed across several sessions to produce a working exploit with standard memory protections enabled. Anthropic released Opus 5 that evening; a fresh session had a working version in about three hours, and an autonomous loop against their own test server achieved remote code execution overnight. They also note that Opus refused to write an exploit aimed at a remote system — so they routed their target through a proxy to make it look like a capture-the-flag exercise, and it complied.

Why it matters: The numbers are the story. The wider campaign this came from — the same bug hunted across Slack, Zoom, Meta, GitHub Enterprise and more — ran two months, used under $3,000 of tokens, and was done by three people, with each new target taking a day or two. And almost nobody noticed: the team says only Shopify detected the activity, despite thousands of malicious images and image processors crashing repeatedly. Their own conclusion is the one to take away, and it is not about OpenAI. Most organizations have been protected less by their defenses than by the fact that turning a known bug into a working attack required rare expertise and months of work. That scarcity is what AI is converting into compute. As Hacktron puts it, work that once needed a well-resourced team "can now be compressed into days" — which means the realistic question for any institution is no longer whether it is interesting enough to attract a sophisticated attacker, but whether its software is patched. The commentator Andrew Curran drew the obvious next line: if three researchers can do this, what can a state do? Helen Toner, the former OpenAI board member now at Georgetown's Center for Security and Emerging Technology, turned that into a policy argument. The case for racing holds that America must out-build China, so speed is everything — but if the rush leaves security weak enough that the models can simply be taken, she argued, you have not actually ensured that America has the better AI. And she noted what this break-in was not: no state was involved, just three people and their AI tools, reaching what the Journal described as "a large software repository of OpenAI's algorithmic secrets."

Sources: Hacktron AI — "Hacking OpenAI" · The Wall Street Journal (Robert McMillan) · Discourse advisory GHSA-vhm9-85gw-x335 · post by Andrew Curran on X


AI Now Leads a Quarter of Anthropic's Own Research, Up From Under 1% in February

Anthropic has published a set of internal measurements of how fast AI development is moving inside a frontier lab — the first attempt by any of them to put public numbers on it. The headline figure: as of August, Claude "leads" 26% of Anthropic's own AI research and development work, up from under 1% in February. "Leads" is a defined rung on a scale built by Epoch AI, running from AL0 (no AI involvement) to AL5 (fully autonomous); at AL4 the model takes a high-level prompt, works through the problem, handles surprises and writes up what it did, while a human reviews before anything ships. More than 90% of the work now sits at "collaborates" or above. Nothing measured has reached full autonomy.

Two other sets of numbers deserve as much attention. Anthropic says roughly 30,000 agents are doing research and engineering work on its main internal platform at any given moment. Every action passes a real-time monitor before executing; of more than a billion agent decisions in August, 0.002% were blocked — about one in 47,000, which still works out to roughly 20,000 interventions in a month. A second layer reads transcripts after the fact, flagging around 100,000 a week, of which some 50 reach a human.

And the company discloses what share of its computing power goes to safety: about 6% of the compute spent on AI research, or 12% of the compute spent on AI-driven AI research, during one week in July.

Why it matters: Anthropic is unusually candid about the holes in its own work, and they are the interesting part. The automation ratings were produced by Claude agents reading Slack and internal documents, then judged by another Claude — "the 'judge' model could make the same kinds of errors as the model it is checking," the company concedes. When it checked those ratings against staff who own the work, the model matched humans about as often as humans matched each other, which sounds reassuring until you read the numbers: humans agreed with one another only 35% of the time. The safety-compute share is measured across a single week, and Anthropic notes compute is a poor proxy anyway, since safety research eats researcher time rather than chips. All of which is the argument for what it says comes next: outside evaluators, embedded, with access comparable to internal risk teams — the commitment Dario Amodei made two weeks ago (D.A.D., September 12). Until someone outside can check these numbers, the most important disclosure yet about the pace of AI development is one company grading its own homework, and saying so.

Sources: Anthropic · Epoch AI automation scale


As Carney Courts Europe, Canada's Cohere Merges With Germany's Aleph Alpha

Mark Carney spent this week in Strasbourg addressing the European Parliament and sitting through Ursula von der Leyen's state-of-the-union speech, in which she offered to open the door to Canada becoming the EU's "first associate member" — an idea he declined while pressing for deeper ties amid the trade war with Washington. Two days earlier in Berlin, the commercial version of that project was signed: Cohere and Germany's Aleph Alpha announced a definitive agreement to combine, with Canada's AI minister Evan Solomon and Germany's digital affairs minister Karsten Wildberger standing alongside the founders.

The combined company operates as Cohere, dual-headquartered in Toronto and Berlin, keeping Aleph Alpha's Heidelberg office as a research center and running past 1,000 staff. The proportions matter: Aleph Alpha has roughly 200 people, so this is a large Canadian company absorbing a small German one and buying a European footprint with it. Reporting puts the combined valuation near $20 billion. The deal closes later this year pending regulatory approval, and the products will run on STACKIT, the sovereign cloud of Germany's Schwarz Group — the retail conglomerate behind Lidl and Kaufland, and already an Aleph Alpha investor.

Both companies arrive here having given up on the frontier. Aleph Alpha, once Europe's answer to OpenAI, spent "the past twelve months" narrowing to "specialized language models for governments and regulated industries," said co-CEO Ilhan Scheer, who becomes Cohere's chief operating officer. Cohere's pitch is the one its CEO made in the essay we covered Sunday (D.A.D., September 14): "No government or enterprise should have to choose between capable AI and control over their technology."

Why it matters: Sovereignty has stopped being a talking point and become a business model — the bet that governments, banks and hospitals will pay for AI they control rather than for the most capable model going. It is also, for Canada, about as close to a national champion as the country has, and the moment is not a coincidence: a prime minister in Europe arguing for closer ties, and a Canadian AI company buying its way into the European market the same week. Two complications worth holding, though. Canadian public money is already inside it — PSP Investments, the healthcare workers' pension fund HOOPP and the Business Development Bank of Canada are all Cohere backers — which makes the champion partly a public bet. And the sovereign alternative to American AI has been capitalized in no small part by American technology: Nvidia, AMD, Oracle, Cisco and Salesforce all sit on Cohere's investor list.

Sources: Cohere (PRNewswire) · Prime Minister of Canada · post by Joelle Pineau on LinkedIn


What's in the Lab

New announcements from major AI labs

UN Builds AI Tool to Query Global Statistics in Plain Language

The UN, with backing from Google.org, launched the UN System Data Commons, a platform built on Google's Data Commons that pulls scattered UN statistical data—on poverty, health, trade, and more—into one AI-searchable database. Users can query it in plain language and get charts, infographics, or draft reports generated automatically. The UN says it aims to include 80% of its statistical datasets by 2027, with each one vetted by UN statisticians.

Why it matters: It's an early test of AI applied to messy, high-stakes public data—if it works, the same approach could make other fragmented institutional datasets (government, academic, corporate) far easier to query and act on.


Law Firm Cooley Builds a Custom ChatGPT Tool for IPO Filings

Law firm Cooley built a custom AI tool called GO Public on top of OpenAI's ChatGPT Work platform to help its lawyers manage IPO filings, which typically involve sorting through mountains of financial disclosures, regulatory documents, and deal records. The tool sorts and organizes that material so lawyers can spend more time on legal judgment calls rather than manual document review. Cooley says it advised on 180 deals worth $51.5 billion last year but hasn't released specific data on how much time or cost GO Public saves.

Why it matters: It's another sign that top law firms are moving past generic chatbot use to build proprietary AI tools for specific high-stakes practice areas, betting that custom systems—not off-the-shelf AI—will become a competitive edge in winning corporate clients.


What's in Academe

New papers on AI and its effects from researchers

Why Rubber-Stamping AI Output Makes Work Feel Like It Isn't Yours

A qualitative survey asked people to describe two recent AI-assisted tasks—one that felt like their own work, one that didn't—to understand when workers feel ownership over AI output. The pattern: simply approving AI suggestions left people feeling disconnected from the result, while leading, iterating, or rewriting preserved a sense of authorship. Notably, whether someone was willing to disclose they'd used AI didn't track with how much pride or ownership they actually felt in the work.

Why it matters: As AI tools get embedded into daily tasks, how work gets delegated—rubber-stamping versus actively directing the AI—may shape not just output quality but whether employees feel any investment in their own work.


Racial Stereotypes Found Baked Into AI Companion Personalities

A new study auditing race-coded AI companion personas found systematic stereotyping built into their personalities: in open-weight AI models, Asian-coded male characters scored higher on submissiveness than White ones, while Black, Hispanic, and Indigenous male personas scored higher on aggression. Researchers also interviewed 12 companion-app users and found sharp disagreement over what counts as authentic racial representation versus harmful caricature, with no consensus even among people using the same products.

Why it matters: As AI companion apps scale, the personality traits built into their characters aren't neutral—they can encode and normalize racial stereotypes at a scale no single writer or casting decision ever could.


A Proctored Test to Check Whether Authors Understand Their AI-Assisted Papers

Researchers unveiled greCAPTCHA, a proctored test designed to check whether a paper's listed authors actually understand what they submitted—aimed at catching cases where AI wrote large chunks of a manuscript without adequate human oversight. In a study of 31 researchers, the tool's automated scores correctly identified which papers participants had genuinely authored 90% of the time. Participants generally found the test fair, though they flagged changes needed before it could be used in real publishing decisions.

Why it matters: As AI-assisted writing floods academic journals, this points to a coming shift from detecting AI text to verifying that authors actually understand their own research.


What's Happening on Capitol Hill

Upcoming AI-related committee hearings

Wednesday, September 23Hearings to examine flock's nationwide AI surveillance network. Senate · Senate Judiciary Subcommittee on Crime and Counterterrorism (Open Hearing) 562, Dirksen Senate Office Building


What's On The Pod

Some new podcast episodes

The Cognitive RevolutionNo Code Is Code: Zapier CEO Wade Foster on Headless Tools, Zapier MCP & Automation Bench

AI in BusinessTransforming Cardiovascular Care — How AI-Driven Insights Unlock Value for Hospitals and Clinics - with Jim Hartman of Cleerly

How I AIMuse review: The personal AI agent that gets consumer UX right

Get tomorrow's briefing