AI This Week
Google DeepMind co-founder and CEO Demis Hassabis is calling on the U.S. to establish a new AI watchdog with the power to screen the world's most advanced models — and coordinate an industry-wide slowdown if dangers mount. The Nobel laureate behind Gemini laid out the plan in a personal manifesto, "A Framework for Frontier AI and the Dawning of a New Age." Hassabis is proposing an AI standards body modeled on FINRA, the private, industry-funded watchdog that polices Wall Street under SEC oversight. Frontier labs would initially voluntarily share models with the body for review up to 30 days before release, before becoming mandatory for U.S. market deployment. Today's AI-driven cyber risks are "warning shots" — within 18 months, biological and nuclear threats could live inside open-source models beyond any government's control. The Trump administration's improvised crackdown on Anthropic's Mythos and Fable models was "a bit of a wake-up call" — Anthropic spent 2½ weeks negotiating model releases with no established rules or playbook. Hassabis wants the new body operational "before year-end." He believes AGI is "probably only a few short years away."
PrismML, a Khosla Ventures-backed spinout from the California Institute of Technology, publicly released compressed versions of Alibaba's open-source Qwen model on Tuesday. The release is sending shockwaves through the AI industry. The company reduced the model from roughly 54 GB to less than 4 GB, allowing all 27 billion of its parameters to run on an iPhone 15 or newer. PrismML CEO Babak Hassibi told CNBC that Apple and other companies have been evaluating the startup's models and measuring their speed, energy efficiency, and performance on devices. PrismML shrinks AI models by drastically simplifying how their internal information is stored — reducing each value from 16 bits to just one or three possible values — significantly cutting the memory required. The compressed models use between 10 and 15 times less memory, generate responses six to eight times faster, and consume three to six times less energy than conventional versions. Larger models running directly on iPhones would allow for more Apple Intelligence features to run on device instead of on Apple's Private Cloud Compute servers, which could reduce Apple's costs and further enhance user privacy. Analysts, however, urge caution. PrismML's claims still need to be proven outside controlled demonstrations, with performance on lengthy prompts, battery consumption, and reliability at scale all critical factors.
Apple is suing OpenAI in federal court in Northern California, alleging trade secret theft — claiming the AI lab took the iPhone maker's intellectual property to develop its own consumer hardware. It's a stunning reversal for two companies that entered a high-profile partnership in 2024, when ChatGPT was integrated into the iPhone's operating system. Apple alleges the misconduct was directed by OpenAI's senior leadership, including Chief Hardware Officer Tang Tan, who is accused of using Apple's confidential project code names during recruiting, asking job candidates to bring Apple hardware components to interviews, and coaching departing Apple employees on how to evade security procedures. The lawsuit also alleges former Apple engineer Chang Liu kept a work-issued laptop and discovered a bug allowing him to access Apple's cloud storage after leaving — and celebrated the exploit. IO Products is also named in the lawsuit. Apple says over 400 former employees now work at OpenAI. Mounting legal woes present another risk to OpenAI as it gears up for what's expected to be a historic IPO.
AI chip startup Etched has secured a $5 billion post-money valuation following a $500 million funding round led by Stripes, with participation from Peter Thiel and Ribbit Capital. The company has booked $1 billion in forward contract orders for its custom "Sohu" inference systems, signaling massive market demand for alternatives to Nvidia's general-purpose GPUs. Instead of building a chip that can handle everything from gaming to scientific simulation to AI training, Etched built Sohu as a custom ASIC designed to do exactly one thing: run transformer model inference as fast as physically possible. The company claims a single 8-chip Sohu server can process around 500,000 tokens per second running Meta's Llama 70B model — outperforming 160 Nvidia H100 GPUs while using less power and taking up less physical space. Etched was co-founded in 2022 by Harvard dropouts Gavin Uberti and Robert Wachen, who previously served as Thiel Fellows. Back in 2023, they struggled to get investors interested — even with a 30-page memo arguing that AI would eventually need specialized chips. Every major investor they pitched passed, and the company was reportedly operating month-to-month, close to running out of cash. Investors now include Jane Street, Hudson River Trading, Jump Trading, Two Sigma, and AI luminaries including Andrej Karpathy, Geoffrey Hinton, and Fei-Fei Li.
A new report from the Center for AI Safety (CAIS) reveals AI agents can now complete 16.1% of real freelance projects at professional quality — a stunning leap from just 2.5% less than a year ago. The frontier has more than quadrupled in under eight months, a concrete signal of how quickly economically capable AI agents are advancing. The Remote Labor Index (RLI) measures how often AI agents can complete real, economically valuable freelance projects — including 3D & CAD, architecture, graphic design, video and animation, audio, data analysis, and web apps — at a quality a paying client would actually accept. Every deliverable is judged by human evaluators against a gold-standard deliverable produced by a paid professional, with the headline metric being the share of projects where the AI's work is judged as good as or better than the human's. Fable 5 reaches the highest automation rate measured so far at 16.1%, roughly double Opus 4.8 at 8.3%, while GPT-5.5 reaches 6.3%.
Anthropic has made a bold move into the pharmaceutical world. The AI company is launching an internal drug discovery program as part of a broader effort to develop artificial intelligence tools for drugmakers. At an event for pharmaceutical executives, biotech founders, and researchers, Anthropic announced Claude Science, a major new product intended to support scientific research. Like Claude Code, Claude Science can autonomously carry out meaningful work when given concise, high-level instructions, with access to tools useful for computational biology and drug development. The platform integrates over 60 scientific databases and computation tools in a single workplace. Anthropic's Head of Life Sciences Eric Kauderer-Abrams said the company plans to prioritize discovering drugs for "neglected" diseases outside the scope of traditional pharmaceutical companies. Anthropic has already inked a collaboration with Bristol Myers Squibb to roll out Claude to over 30,000 employees.
OpenAI has introduced GeneBench-Pro, a research benchmark designed to assess whether AI agents can perform the complex, judgment-intensive analysis required in real-world computational biology. Unlike conventional benchmarks focused on factual recall, GeneBench-Pro measures what OpenAI calls "research taste" — the sequence of judgment calls in scientific analysis, from interpreting ambiguous data to deciding whether findings are robust enough to inform downstream research. The benchmark comprises 129 problems spanning ten domains, including statistical genetics, cancer genomics, clinical diagnostics, and pharmacogenomics. Reviewers estimated each task would require 20 to 40 hours of work by a human expert — at an estimated $200 per hour — compared with AI inference costs of just a few dollars per task. OpenAI's GPT-5.6 Sol model achieved a pass rate of just 28.7%, rising to 31.5% in Pro mode — still a significant leap from GPT-5's sub-5% score on the original GeneBench. To encourage independent evaluation, OpenAI is open-sourcing ten tasks on Hugging Face and providing a 50-question subset to Artificial Analysis for third-party benchmarking.
AWS announced it is investing $1 billion in a new Forward Deployed Engineering unit that will help its customers build and deploy AI systems. The organization will embed teams of five to six engineers directly inside customer environments for roughly 45-day engagements, building and shipping production agentic AI systems alongside clients' own staff. Speed is the driving force. Human engineers oversee AI agents that handle the software writing and system deployment work — shrinking what normally takes months down to days. The new unit will be seeded with "thousands" of FDEs, said Francessca Vasquez, AWS' vice president of frontier AI engineering and services. Named early customers include the Allen Institute, Cox Automotive, the NBA, the NFL, Ricoh, and Southwest Airlines. Both OpenAI and Anthropic have launched their own FDE joint ventures in recent months, valued at $4 billion and $1.5 billion, respectively. AWS funds this one entirely from its own balance sheet with no outside investors.
Shares of Meta closed up nearly 9% on Wednesday following news that the company is building out a new cloud business — one that could finally justify its staggering infrastructure bets. In 2024, Meta's capex totaled $37.2 billion, before rising to $69.6 billion last year — and is projected to nearly double this year to $135 billion at the midpoint of its guidance range. In the past, Meta defended its AI investments by saying they improved its advertising business, but the spending began to severely crimp free cash flow, making some investors uncomfortable. The company is debating whether to offer access to AI models hosted on its infrastructure or sell access to raw computing power. The move throws Meta into fierce competition with Amazon, Microsoft, Google, and CoreWeave — and shares of neocloud companies CoreWeave and Nebius Group both plunged about 12% each on the news. Of the four U.S. hyperscalers, Meta is the only one that doesn't currently sell cloud infrastructure and services.
OpenAI is in preliminary talks to hand the U.S. government a 5% stake in the company — a slice worth roughly $42.6 billion. The figure is based on the AI lab's record-breaking March funding round at a post-money valuation of $852 billion. CEO Sam Altman argued that giving the public a financial interest in OpenAI is the best way to share the upside of AI, according to two people familiar with the talks. The proposed arrangement envisions other U.S. AI companies — such as Anthropic, Google, and Meta — ceding similar stakes to the government through a sovereign wealth fund vehicle. Altman and other OpenAI executives suggested allotting 5% of equity to a vehicle similar to the Alaska Permanent Fund. Any deal might require an act of Congress to implement. Altman has been in active talks with Trump, Commerce Secretary Howard Lutnick, and Treasury Secretary Scott Bessent. Senator Bernie Sanders recently called for part-nationalization of the AI industry, with dividends going to the public. It is not clear whether any companies would agree to OpenAI's proposal.