July 24, 2026

D.A.D. today covers 7 stories — about a 6-minute read. What's New, What's Innovative, What's Controversial, What's in the Lab, and What's in Academe.

The Daily AI Digest is a daily AI briefing automated by Alexander Panetta — a veteran political journalist tracking the field during a Master's in AI Management at Georgetown University.

D.A.D. Joke of the Day: I asked AI to help me draft my resignation. It kept telling me to "regenerate the response" — which is exactly what my career needed.

What's New

AI developments from the last 24 hours

Microsoft Builds Its Own Cheap Models to Lean Less on OpenAI

Microsoft is making a quiet but consequential bet: that for most everyday AI tasks, you don't need a giant frontier model at all. In a post from its AI division, the company said it has built small, specialized models—its "MAI" family—that now match or beat frontier models on common work inside GitHub Copilot and Excel while using far fewer resources. The coding model, MAI-Code-1-Flash, already serves millions of developers and, Microsoft says, posts a roughly 10% higher code-acceptance rate than comparable small models from OpenAI and Anthropic (GPT-5.4 Mini and Claude Haiku 4.5) while burning about 10% fewer tokens. It then retrained that same model for spreadsheets, producing an Excel assistant it rates "on par with GPT-5.6 for the most common tasks"—but cheap enough to run on older A100 chips rather than only the latest, scarce accelerators. Microsoft calls the method "hill-climbing": because it controls the whole stack—the model, the agent harness around it, and the product's own evaluations—it can keep tuning a small model against the exact tasks users perform until it's good enough to replace a bigger one. CEO Satya Nadella framed the stakes in an essay titled "Frontier Diffusion & Control": "In a world where software has real marginal cost for the first time"—his point that AI, unlike traditional software, costs real money every time it runs—the winning move is to "take saturated frontier capabilities and deliver them at scale at lower cost," reserving frontier models "for frontier needs." Tellingly, Microsoft says it is now routing traffic across its own products to MAI "whenever our models match or outperform frontier alternatives"—including those from OpenAI and Anthropic—and extending the approach to Outlook, Copilot Chat, and PowerPoint.

Why it matters: Two shifts hide inside a dry engineering post. The first is economic: AI has given software a per-use cost for the first time, turning "which model" into a line-item decision. Microsoft's answer—small models specialized on your actual workflows—says the frontier is becoming a commodity for routine work, and the money is in optimizing cost per outcome, not chasing the biggest model. That's the same logic driving this week's fight over cheap Chinese open models (D.A.D., July 23), now coming from inside the world's largest software company. The second is control: Microsoft is OpenAI's biggest backer, yet here it is building its own models and pointing its products at them, deliberately keeping the "harness, memory, and skills" outside any single model so it can swap providers at will. Nadella's word for it—"Control"—marks how far Microsoft has moved to reduce its dependence on OpenAI. For enterprises, the takeaway is the template Microsoft is now selling through its Foundry platform: you probably don't need to pay frontier prices for most agentic work; a smaller model trained on your own tasks may do it for a fraction of the cost.

Sources: Microsoft AI — Hill-climbing MAI models · Satya Nadella (@satyanadella)


What's Controversial

Stories sparking genuine backlash, policy fights, or heated disagreement in the AI community

Washington's Own Tests Find China's Kimi Trails on Hacking—and the Doves Pounce

The Trump administration's own evaluators just poured cold water on the "Kimi panic" driving this week's crackdown. In a joint assessment published Wednesday, the U.S. Center for AI Standards and Innovation (CAISI, the renamed federal AI-testing body) and the U.K.'s AI Security Institute found that Moonshot's Kimi K3—the Chinese model at the center of the distillation fight (D.A.D., July 23)—performs "significantly below" the leading U.S. frontier models on cyberattack tasks. On an exploit-writing benchmark built by Carnegie Mellon, Kimi never once reached the most dangerous outcome, full "arbitrary code execution" (0 of 41 tests), where top U.S. models succeeded on 20. Turned loose on a simulated 32-step corporate-network breach, Kimi got about halfway (step 17) versus 28.5 for the best U.S. models. But the report is not all reassurance: Kimi K3 is now the most cyber-capable open-weight model on earth, beating China's previous best (GLM-5.2); its safeguards did nothing to stop it from attempting offensive hacking when asked; and in one of ten tries it autonomously completed the simulated attack—showing it can breach a weak, undefended system on command. One caveat cuts against the headline finding: the U.S. models were tested with their safety guardrails switched off to measure raw capability, while Kimi was tested as shipped—and simply didn't refuse.

The reactions split along the fault line D.A.D. has tracked all week. Commerce Secretary Howard Lutnick seized on the report—"Kimi K3 remains behind America's leading frontier AI models," he wrote; "the United States continues to lead." David Sacks—the venture capitalist and former White House AI adviser—went further: "The Kimi Panic needs to stop… Let our horses run," he posted, arguing American models stay ahead as long as Washington doesn't "sabotage ourselves with unnecessary rules"—a pointed jab, from the crackdown's most prominent critic, at the sanctions and export threats coming from Treasury and the science office.

Why it matters: The timing is everything. For a week the hawks—Treasury's Bessent, the science office's Kratsios—have cast Kimi as a national-security menace to justify sanctions and a possible ban. Now the government's own labs have measured it and concluded it trails U.S. models on exactly the capability, cyberattack, most often invoked as the threat—handing Commerce and outside allies like Sacks the data to argue the panic is overblown and restrictions would only hobble American firms. That's the split we flagged—Commerce against the Treasury-and-science-office hawks—now fought with evidence instead of tweets. But the report is a Rorschach test: the same findings show Kimi is the best open model yet at hacking, will do it without complaint, and can autonomously breach a weak network—precisely the "open weights have no brakes" danger the restriction camp warns about. Both sides can, and will, claim it. What's no longer in doubt is that the fight over Chinese open models is now anchored to real measurements—and that the number each camp cites will depend on which risk it fears most.

Sources: CAISI/NIST — UK AISI/CAISI assessment · Howard Lutnick (@howardlutnick) · David Sacks (@DavidSacks) · U.S. Dept. of Commerce (@CommerceGov)


What's in the Lab

New announcements from major AI labs

ChatGPT Can Now Pull Your Health Records Into Any Conversation

OpenAI is rolling out a Health feature in ChatGPT for logged-in U.S. users 18 and up, letting people connect Apple Health data and medical records so the chatbot can reference that information across any conversation, not just a dedicated health tab. OpenAI says over 300 million people already ask ChatGPT health questions weekly, and that most of those conversations happen outside a standalone health space—hence the broader rollout. The company says connected health data won't be used to train its models or for ad targeting.

Why it matters: As ChatGPT becomes a default place people take health questions, this makes it more useful for everyday medical queries—but it also puts sensitive records into a general-purpose chatbot, so OpenAI's privacy assurances are worth reading before you connect anything.


What's in Academe

New papers on AI and its effects from researchers

Researchers Turn Medical Case Studies Into Interactive Training Games

Researchers built MedGame, a system that converts static medical case studies into interactive branching-story games, along with a benchmark of 5,000 clinical cases to test how well AI models generate these narratives. The team reports that fine-tuning smaller, open-source models on this task closed much of the gap with commercial GPT-4-class systems. A small pilot found students rated the game-based cases as more engaging and useful than traditional text materials, though no hard numbers were given.

Why it matters: It's an early example of AI reshaping how professional training gets built and delivered—not just what trainees study, but how institutions teach judgment under uncertainty.


Abuse Victims Get Poor Help From Search, Reddit, and AI Chatbots

A new study tested how well web search, Reddit forums, and AI chatbots respond to victims of tech-facilitated abuse—stalking via GPS trackers, hidden cameras, or spyware—using a decade of real victim queries. The surprise: purpose-built survivor-support chatbots performed worse than general-purpose assistants like ChatGPT on most measures. More than 65% of victim search queries surfaced potentially malicious links, over 20% of Reddit responses were toxic, and most conversational AI systems failed to offer risk-aware guidance or point to concrete help resources.

Why it matters: As abuse increasingly involves everyday tech, the tools victims turn to for help—including AI chatbots marketed as safe—are largely unequipped to recognize danger or guide them to real support.


ChatGPT Access Didn't Inflate College Grades, Data Shows

A study tracking 156,000 students and nearly 88,000 courses at a large U.S. university from 2015 to 2025 found no evidence that ChatGPT's availability inflated grades. Researchers compared courses before and after the chatbot's release, sorting them by how vulnerable assignments were to AI use (take-home essays versus in-class exams). Grades and self-reported understanding showed no significant difference, even among previously lower-performing students—undercutting the assumption that easy AI access lets students coast to better marks.

Why it matters: As employers and schools debate whether AI is quietly eroding academic rigor, this data suggests the more alarming fears about mass grade inflation haven't materialized, at least so far.


Stress-Sensing Wellbeing Apps Carry Hidden Fairness Risks

A study based on interviews with 14 researchers and practitioners across five countries examined passive sensing systems—apps and wearables that infer depression, stress, or cognitive load from phone and sensor data. The researchers found fairness problems don't just stem from demographics like race or age; factors like how comfortable someone is with being monitored, or how irregular their daily routine is, can also skew results. The study catalogs 15 distinct fairness risks that emerge across building and deploying these systems, along with ways to address them.

Why it matters: As employers, insurers, and health apps increasingly use passive sensor data to assess wellbeing, this research flags bias sources that standard fairness audits—focused on demographic categories—may miss entirely.


What's Happening on Capitol Hill

Upcoming AI-related committee hearings

Friday, July 24Building an AI-Ready America: How AI Is Creating Opportunities Across America's Workforce House · House Education and Workforce (Hearing)


Wednesday, July 29Hearings to examine the impact of AI on the workplace. Senate · Senate Health, Education, Labor, and Pensions Subcommittee on Employment and Workplace Safety (Open Hearing) 430, Dirksen Senate Office Building


Wednesday, July 29Hearings to examine the AI deception machine, focusing on deepfakes, chatbots, and the new frontier of senior fraud. Senate · Senate Aging (Special) (Open Hearing) 562, Dirksen Senate Office Building


Thursday, July 30Hearings to examine intelligent networks, focusing on powering artificial intelligence and transforming communications. Senate · Senate Commerce, Science, and Transportation Subcommittee on Telecommunications and Media (Open Hearing) 253, Russell Senate Office Building


What's On The Pod

Some new podcast episodes

AI in BusinessAgentic CRM for SMB Automation - with Sharif Karmally of Salesforce

How I AIComputer & browser use in Codex (5 real examples)

Get tomorrow's briefing