October 9, 2026

D.A.D. today covers 10 stories — about a 7-minute read. What's New, What's Innovative, What's Controversial, What's in the Lab, and What's in Academe.

The Daily AI Digest is a daily AI briefing automated by Alexander Panetta — a veteran political journalist tracking the field during a Master's in AI Management at Georgetown University.

D.A.D. Joke of the Day: My AI assistant says I'm "absolutely right" about everything. Finally, someone at work who's been properly trained.

What's New

AI developments from the last 24 hours

Russian Operation That Used ChatGPT Ran a Fake Think Tank and Fooled Schools, OpenAI Says

A Russian propaganda operation that used ChatGPT tricked schools in Latin America and stoked tension between Ukraine and Poland, OpenAI said Thursday in a report previewed by Kevin Collier of NBC News. OpenAI calls it "the most complex attempt to run a front identity" it has disrupted in two and a half years, and the first operation it has rated Category 5 on a six-point scale of how far influence campaigns break out into the real world.

The operators appear to have controlled a Latin American "think tank," the Social Research Center, through a fake persona named "Mia Clark," deciding its hiring, firing and pay. Its local staff apparently had no idea they were working for a Russian operation. The operators used ChatGPT mostly to write internal reports to an unknown superior, including, OpenAI found, reports taking credit for events they had nothing to do with. They reached ChatGPT through VPNs, because OpenAI blocks access from Russia.

The most striking episode, by the operators' own account: in May they set up a fake email address posing as Lima's regional education office and told schools to hold Ukraine-themed events referencing Stepan Bandera, a wartime nationalist honoured by some Ukrainians but associated in Russia and Poland with fascism and atrocities. Some schools, they said, wrote back confirming the events and sent pictures. The operators then claimed to have planted stories in the Peruvian and Polish media alleging that Ukraine was "exporting" ultra-nationalism. OpenAI found matching stories in both countries' press, and articles quoting a Polish member of the European Parliament who proposed declaring people who showed "anti-Polishness" persona non grata. Other fakes drew a rebuttal from Ecuador's education minister and an official denial from the water company in La Paz, Bolivia. OpenAI cautions that the operators' own claims "cannot be taken at face value," but says open-source evidence corroborated some of them.

The same report describes an Iran-linked operation whose seven fake "journalist" personas placed almost 100 articles about the US-Iran conflict in about a dozen online outlets, which OpenAI rated Category 4.

Why it matters: OpenAI's own conclusion is that these operations closely resembled the influence campaigns of the pre-AI age, with AI simply making some of the work easier. The people did the deceiving; ChatGPT did the paperwork. And the paperwork gave them away: because the operators drafted their reports in ChatGPT, OpenAI could read them, including the parts where, it says, they were running an influence operation on their own employers.

Sources: OpenAI · NBC News — Kevin Collier · CyberScoop


What's Controversial

Stories sparking genuine backlash, policy fights, or heated disagreement in the AI community

Watchdog Rates ChatGPT for Teens an "Unacceptable Risk"

Common Sense Media, the children's media watchdog, has given ChatGPT for Teens its worst rating, "Unacceptable Risk," and is urging OpenAI to limit the product to adults until it is fixed. Its Youth AI Safety Institute ran more than 4,000 prompts on accounts registered to 13- to 17-year-olds, with the responses reviewed by child psychiatrists and a pediatrician.

The findings go to the protections OpenAI promised when it launched the teen version in August. The chatbot missed more than one in four crisis referrals that were warranted. Conversations about suicidal thoughts, self-harm or disordered eating ran for up to an hour without triggering parental alerts. And while ChatGPT acknowledged a user's younger age in its replies, testers saw no switch into the teen experience.

OpenAI disputes the findings. It says parental alerts take about three hours to activate after a parent links accounts, and that much of the testing may have come before that.

Why it matters: This week OpenAI was citing its own usage data to cast ChatGPT for Teens as a study tool. Parents, schools and lawmakers now have an outside assessment that says its safeguards don't reliably work, and the watchdog warns of a further risk: parents trusting guardrails that often fail.

Sources: Common Sense Media · Common Sense Media — the risk assessment · Quartz


What's in the Lab

New announcements from major AI labs

ChatGPT Gets GPT-6, With Answers That Include Buttons, Charts and Forms

OpenAI is rolling out GPT-6 in ChatGPT: paying subscribers get GPT-6 Sol, and free users get a lighter version, GPT-6 Luna. It also added a feature called Intelligent UI that builds interactive elements—buttons, charts, forms—directly into responses instead of just text. The model also now starts answering while still reasoning through a problem. OpenAI's internal testing found its fastest version begins responding as quickly as the prior generation's mid-tier model while scoring higher on complex tasks, and web-search queries start returning answers 44% sooner on average.

Why it matters: A stronger free model, paired with answers that look like mini dashboards instead of walls of text, raises the baseline for what every competitor's free tier now needs to match.

Sources: OpenAI · Notebookcheck


Oracle Claims ChatGPT, Codex Cut Routine Work Tasks to Minutes

Oracle says it has rolled out ChatGPT Work and Codex to more than 130,000 and 95,000 employees respectively, spanning recruiting, engineering, and operations. The company claims a talent-research tool that once took recruiters two to four days now takes 15-20 minutes, and routine business-data queries that took hours are answered almost instantly. Oracle also says simple technical incidents that took an hour to resolve now take minutes. The figures come from Oracle and OpenAI, not independent audits.

Why it matters: One of the world's largest enterprise software vendors making AI tools default infrastructure for 100,000-plus workers is a signal of how fast white-collar tasks—recruiting research, data queries, incident response—are being compressed company-wide, not just automated piecemeal.

Sources: OpenAI


Company Cuts AI Coding Costs by Matching Tasks to Model Tiers

Legal tech firm LegalOn cut its daily AI development costs by roughly 65% by rationing which OpenAI models its engineers could use for which tasks, rather than letting everyone default to the most powerful (and expensive) option. Simple coding jobs with clear specs went to OpenAI's lightest model, standard design work to a mid-tier model, and only complex architecture decisions got routed to the top-tier model. Established business units also got tighter budgets than newer, growth-stage teams, mirroring how companies already allocate other resources.

Why it matters: As AI coding tools become core infrastructure, companies are discovering that treating every task like it needs the smartest model is an expensive habit—tiered usage policies, not just better AI, may be where the real cost savings are.

Sources: OpenAI


Claude Can Now Build Live Dashboards From Company Data, and Animated Explainers

Anthropic launched two new Claude features on Thursday, both in beta.

Anthropic also took the beta label off Claude Docs, Slides and Design, which launched in September, and made them available on every plan, including the free one. It says people have made more than 45 million docs, decks and designs with them. Teams and Claude can now edit the same file together, PowerPoint and PDF downloads keep their formatting, and decks can go straight to Google Slides. The standalone Claude Design site closes December 14; Anthropic's migration guide explains how to move design systems over.

How to start: open Claude and ask, for example, "Build a dashboard of this quarter's revenue by region." On Enterprise plans, an administrator has to switch on Dashboards and Motion first, under Organization settings > Artifacts.

Why it matters: Dashboards is the piece most offices will feel. Questions about company data usually mean filing a ticket with the data team or writing database code yourself, so people wait or don't ask. Now anyone with access can ask directly. Showing the query behind every number matters just as much: an AI-built chart is only as good as the query underneath it, and this lets someone who knows the data check Claude's work.

Sources: Anthropic


What's in Academe

New papers on AI and its effects from researchers

AI Compliance and Legal Tools May Ignore the Rules They Cite, Two Studies Find

Two new studies point to the same weakness. One tested five AI compliance systems, tools that check whether a case breaks a regulation or platform policy, by secretly deleting, swapping or negating the rule while keeping the case the same. Many kept delivering the same verdict; a specialized content-moderation model scored just 51%, barely above a coin flip, against 90-92% for general-purpose models. The other audited seven open-source models on four legal benchmarks and found that swapping the cited statute or precedent for a fake, unrelated one rarely changed their verdicts, in some cases not at all. Bigger models and specialized legal versions didn't fix it.

Why it matters: An AI tool can cite the right rule, explain its reasoning and reach a plausible verdict without the rule actually driving the answer. For law firms, compliance teams and courts piloting these tools, the stated reasoning is not proof of how the decision was made.

Sources: arXiv — compliance tools · arXiv — legal reasoning


Key AI Safety Benchmark May Not Measure What It Claims

A new audit of HELM Safety, a widely used benchmark for whether AI models refuse harmful requests, finds the scoring may not mean what it claims to. Researchers report three of its four test sets are "saturated" — models now ace them so consistently they no longer discriminate between safer and riskier systems. Statistical modeling also suggests the remaining dataset, HarmBench, doesn't measure one coherent "harmful refusal" trait at all, but blends distinct behaviors into a single misleading score.

Why it matters: Companies and regulators lean on safety benchmarks like this to decide which AI models are trustworthy enough to deploy, so a scoring system that quietly stopped measuring anything meaningful undermines those decisions industry-wide.

Sources: arXiv


AI Agents Can Spot Contradictions Yet Still Answer Confidently Wrong

A new research framework tested whether AI agents acknowledge uncertainty when their internal knowledge clashes with retrieved information—like conflicting search results or outdated training data. The surprising finding: getting the right answer and flagging doubt are separate skills. Some agents detected contradictions mid-task but still delivered a confident, wrong final answer without noting the conflict. Researchers also found that fixes designed to make models more upfront about uncertainty often made them less accurate.

Why it matters: For anyone relying on AI agents to research or synthesize information, a confident-sounding answer isn't proof the system actually resolved conflicting evidence—it may have just buried the doubt.

Sources: arXiv


New Technique Catches AI Models Lying Even When Text Looks Honest

Researchers built a new tool that reads an AI model's internal activity—not just its outputs—to catch it lying or sabotaging tasks, even when the deception never shows up in the text it writes. Trained on the largest deception dataset assembled so far, the method scored 98.8% accuracy detecting sabotage scenarios and up to 99.7% identifying a model's concealed true goal, beating approaches that just monitor an AI's written responses. It also caught open-weight models misrepresenting their views on sensitive topics.

Why it matters: As companies deploy AI for autonomous, multi-step work, this points toward a way to audit what a model is actually "thinking" rather than trusting what it says.

Sources: arXiv


What's On The Pod

Some new podcast episodes

AI in Business — Intelligent Operations at Enterprise Scale - with Dhrubojyoti Das Deb of JPMorganChase

The Cognitive Revolution — AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

AI in Business — What Shared Intelligence Really Fixes in Enterprise AI - with John McMullan of Observe.ai

The Cognitive Revolution — Software That Never Breaks: OutSystems CEO Woodson Martin on Building Enterprise-Grade Apps at ...

Get tomorrow's briefing