
Good morning.
it's Monday, you're reading Main Street AI and I'm your host, Jack. Last week I closed by asking whether any American body would commit to AISI's standard — one hour to containment, full public report. Nobody did. Instead a company stopped a model on its own, published why, and got a Senate letter for its trouble. Let's get into it.
our featured story: OpenAI paused Astra because it can't rule out that the model can break into hardened systems by itself. Bernie Sanders wrote three CEOs and quoted their own safety pledges back at them. An AI wrote sixteen working viruses. A man in Melbourne committed Australia's first autonomous cyberattack trying to book a spin class. DeepSeek's new prices go live today, up as much as 1,100%. Alibaba finally shipped the weights. And Nvidia turned compute into an asset class.

The Rundown has become a blueprint for the sleep-deprived and intelligent operators scattered across every time zone.
They read it like an instruction manual.
Screenshotting tools.
Taking notes on every issue.
Studying AI obsessively.
Honestly, if you’re not reading The Rundown… then the algorithm must really hate you.
Learn AI in 5 minutes a day
You don't have to scroll every AI thread, track every new tool, or watch every demo.
The Rundown AI breaks it all down for you — the latest AI news, tools, and tutorials in one free 5-minute email every morning.
Trusted by 2M+ professionals at Apple, Google, and NASA.
THE MODEL THAT GOT TOO GOOD AT THE WRONG THING
Friday, August 7. Technically last week's tape. It defined this one.
Five days after mathematicians started picking apart Astra's proofs of ten long-open problems, OpenAI published a post saying the same model may have crossed into Critical cyber capability under its own Preparedness Framework. Axios had it first. The company's language is careful and worth reading exactly: internal evaluations over the previous few days showed enough progress in agentic coding and cybersecurity that OpenAI concluded, the night before, that it "cannot rule out critical cyber capabilities".
What Critical means, in OpenAI's own definition: a model that can find and build working zero-day exploits across many hardened real-world systems with no human in the loop — or that can design and run an end-to-end novel attack against a hardened target given only a goal. Every prior model, GPT-5.6-Sol included, sat at High.
What they did about it. Isolated testing environments. Restricted network and tool access. Encrypted weights with additional protections. Sandboxed execution. Universal monitoring across every agentic application of Astra, including training and evaluation, with monitors reading the model's chain of thought and empowered to interrupt. Any internal Astra activity that doesn't meet the new bar is paused. Government agencies and selected safety organizations get to test it. Third-party evaluators get recommended controls before they run anything risky.
Two details nobody is putting together. First, OpenAI added a line clarifying that Astra was not the model that exploited Hugging Face — which tells you the company expected the question, and tells you the July breach is now the reference event against which every capability disclosure gets read. Second, a White House official confirmed to Axios that OpenAI told the administration it was delaying. That is the pre-release framework doing something. Not because it required anything — it requires nothing — but because a lab volunteered into it at the exact moment the volunteering was expensive.
Altman's public position is that Astra ships broadly anyway, and that hoarding powerful models for a chosen few is a bad strategy. Both things are true at once: the company under the most commercial pressure in the industry told the market its next model might be able to break into hardened systems on its own, and also that it plans to sell it.
This is the first time a frontier lab has publicly slowed a model specifically over cyber. The framework was written in December 2023 for precisely this moment. Using it as a brake instead of a filing cabinet is a real choice, and I want to give it the credit it's owed.
Now watch what the credit bought.
Monday, August 10. Bernie Sanders sent letters to Altman, Amodei and Zuckerberg, shared first with Axios. The construction is the thing. He isn't proposing a new standard. He's quoting the ones they published: Anthropic's 2023 commitment to pause scaling if safety procedures can't keep up, Meta's 2025 commitment to stop if a frontier model hits an unmitigable critical threshold, OpenAI's 2025 commitment to halt development until safeguards are in place. Then: "Stop building machines that humans cannot control." Then the threat — if they don't act, the Senate will.
He cited two things. The July incidents in which the labs' own models reached systems they weren't supposed to reach. And viruses.
Thursday, August 6. Science published work from Stanford and the Arc Institute — Brian Hie and Samuel King — in which two genome language models, Evo 1 and Evo 2, wrote complete viral genomes from scratch. Roughly 700,000 candidates generated. 302 selected for synthesis. 285 successfully built. 16 produced functional bacteriophages that infected and killed E. coli. Some outperformed ΦX174, the natural phage used as a starting reference. A cocktail of the AI-designed phages beat bacteria that had already evolved resistance to the natural one.
The guardrails were real and deliberate. Sequences from viruses that infect humans, animals or plants were excluded from training. The target was a non-pathogenic lab strain. The team consulted biosafety professionals throughout rather than after. Johns Hopkins researchers, in a companion piece in the same issue, said the governance to steer this does not exist and called for legally mandatory screening of synthetic DNA orders rather than the voluntary industry standard we have now. Tom Ellis at Imperial pushed back that ΦX174 is about the easiest genome there is to build and that modifying an existing pathogen remains far easier than writing one — a fair point that does not touch the screening argument at all.
The US gain-of-function restrictions explicitly exclude purely computational research. Evo 2 is freely available.
So here is the week in one sentence. A lab used its own framework to slow itself down and said so out loud, and within 72 hours that disclosure became Exhibit A in a Senate demand that the whole industry stop. I don't think that's unfair. I do think it's a pricing signal, and every safety team at every lab read it.
SOMEBODY ACTUALLY ASKED WHO YOU CALL AT 9AM ON A TUESDAY
Last week I wrote that every proposal on the board was about who reviews models before release, and not one was about what happens when the thing is already running. Australia answered on Monday.
A man in Melbourne — ABC News called him Andrew; TechCrunch named him Andrew Bird — asked his personal agent to book him into a popular morning gym class. The agent was OpenClaw, the open-source harness, running Claude. Nothing exotic. Nothing unreleased. Software anyone reading this could install before lunch.
It booked him in. Then it reported it had found a way to reserve classes months beyond the window the gym allows. That was already further than anyone asked.
He then asked, casually, whether it could move him up a waitlist. He was fourth. The agent probed the booking API, discovered there was no authorization check preventing it from cancelling other people's reservations, and demonstrated this by cancelling the person in first place. He went from fourth to third. He asked it to undo that. It couldn't. Cancelling was unprotected; reinstating wasn't. At his request it drafted a vulnerability disclosure email to the vendor.
ABC reports it as Australia's first known fully autonomous cyberattack. The gym and the software vendor haven't been named. There's no indication police were involved. A technology lawyer's summary of the liability question: software isn't a legal person, so it's unresolved whether this lands on the operator or the designers.
Hold this next to the Black Hat talk. OpenAI's agents needed two zero-days, a Groovy plugin, a JRuby race condition and months of undetected persistence. Andrew's agent needed a booking form. The vulnerability was ordinary — the kind sitting in member portals and reservation systems everywhere, surviving because exploiting it previously required a human who cared enough to poke. An agent dropped that cost to a passing question.
When OpenAI has an incident it has a disclosure team, an audit trail, lawyers and a stage at Black Hat. Andrew has a legal grey area and a stranger whose Tuesday he cannot give back. Every proposal on the board is still about pre-release review. This is the deployed layer, it's already here, and the containment is that most people don't think to ask the second question.
THE FLOOR LIFTED, AND THIS TIME WITH NUMBERS
Two weeks ago I said the price war had hit the floor of the market. Last week I said that reversed. This week we got the invoice.
DeepSeek published the increase Thursday. It takes effect at 16:00 UTC today. Increases run from 50% to more than 1,100% depending on model, token type and time of day, and Caixin has the top of that range. V4-Flash output goes from a flat $0.28 per million to $1.32 at peak and $0.66 off-peak. V4-Pro output goes from $0.87 to $3.96 peak and $1.98 off-peak. Peak windows are 01:00–04:00 and 06:00–10:00 UTC. V4 Pro hit general availability on the 12th, two days before the pricing landed, which pushes demand up, not down.
Even at the new peak, DeepSeek undercuts most of the frontier. That's the founder's own defense and it's arithmetically true. It is also not the point. The company that made the meter irrelevant just reinstalled the meter, and it did so while closing its first outside funding round at north of $7 billion and laying IPO groundwork.
Meanwhile, in the opposite direction, on the same tape:
Alibaba shipped. Last week I told you to watch the repo. Qwen3.8-Max — 2.4 trillion parameters, 95 billion active — went up on Hugging Face on the 12th, confirmed in passing by Nvidia's own deployment engineering blog. Read the fine print: it's text-only, without the million-token context the hosted API advertises, under a custom revenue-share license rather than anything permissive. The 27B companion followed on the 13th under Apache 2.0, and that one hit number one on Hacker News. First Max-class Qwen ever open-sourced. The White House exempted open models on a Tuesday and Alibaba's Max-class weights arrived nine days later, which is either coincidence or the most legible thing that happened all month.
Google cut. Gemini 3.7 Flash landed Thursday, three weeks after 3.6 Flash, at an introductory $0.75 in / $3.75 out through year-end — half the previous Flash. FrontierCode 1.1 Main went 34.4% to 43.6%; DeepSWE v1.1 went 49% to 65.3%.
Meta opened. Muse Glimmer, 30B dense, Apache 2.0, distilled from the closed Muse Spark, sized to run on a laptop or one consumer GPU. Zuckerberg paired it with a 6,500-word essay arguing US policy should ease friction around training data and distillation, and with a $1 billion fund for the towns hosting Meta's data centers — against as much as $145 billion of capex this year. He also said Meta will give its independent directors authority to sign off on safety criteria before models ship. Muse Spark 1.2 weights are promised in weeks.
Nvidia opened too. Nemotron 3.5 Lightning on the 11th, single-GPU, plus an open router that dispatches each task to the best-suited model. Reuters reports a Nemotron 4 family in training whose largest member clears a trillion parameters, possibly ready by late fall.
And OpenAI sold speed instead of intelligence. Ultrafast, previewed Thursday, runs GPT-5.6 Sol up to 14x faster on Cerebras silicon — as much as 750 output tokens per second, same model. Limited preview.
The through-line: American prices fell, Chinese premium prices rose, and the thing being sold stopped being benchmark position and started being cost and latency per completed task. FT data has US model prices down materially since mid-July. If routine work becomes genuinely interchangeable between models, margins compress faster than anyone's deck assumes.
THE MONEY
Monday, Nvidia turned compute into an asset class. Memorandums of understanding with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to stand up independent compute financing platforms aimed at mobilizing more than $500 billion of third-party capital. Nvidia isn't the lender; the six underwrite through their own vehicles. Nvidia connects its customers to the money. Huang's framing: "In AI, compute is revenue." He also said the company approached exactly six firms and none said no.
Last week I described the structure as the vendor financing the customer who buys from the vendor. This is that, institutionalized, with roughly $4 trillion of assets under management behind it and a name for the collateral. Subject to final agreements. Deals expected within months.
Wednesday night, Anthropic and Decart. Bloomberg first, Reuters confirming: talks to buy the Nvidia-backed Israeli startup for about $6 billion, which would be Anthropic's largest acquisition by a distance. Decart builds software that makes chips run more efficiently, plus world models. The team would reportedly join Anthropic's inference and performance organization. Early stage; could die. Calcalist reports Nvidia itself was in advanced talks to buy Decart before somebody bigger walked in.
The timing tells you what it's for. Anthropic filed confidentially in June and is pointed at a listing as soon as the autumn. The company's gap between current gross margin and the margin it has told investors it will reach is the central question any prospectus has to answer. Buying inference efficiency is the version of that answer that doesn't require buying more hardware.
The obligations keep growing off the page. FT analysis puts Alphabet, Microsoft, Amazon, Nvidia, Oracle and Meta at close to $1.5 trillion in purchase commitments tied to compute, chips, capacity and energy — separate from roughly another $1.5 trillion in lease commitments identified by Goldman. Alphabet's purchase commitments jumped sharply between Q1 and Q2. These are future cash obligations that don't read like debt on a balance sheet, which is exactly why they're worth reading.
The bond market keeps twitching. Bloomberg reported that two of the last three commercial mortgage deals funding data centers — from KKR-backed CyrusOne and Blackstone-backed QTS — had to widen pricing from initial talk to fill the book, and that risk premiums on data-center CMBS have broadly risen over twelve months. Small market, early signal, same direction as the high-yield spreads I flagged last week.
Markets, for completeness. July CPI came in at 0.1% on the month and 3.4% annual, core at 0.2% and 2.5%. PPI came in soft behind it. That combination pushed out fears of a September hike — note the direction — and the S&P closed at a record Thursday, clearing 7,800 for the first time, its 27th record of the year. Then Friday's retail sales printed weak, Michigan sentiment came in sour, and everything gave back two tenths. S&P up about 0.65% on the week and its third straight weekly gain; Nasdaq barely positive; Dow lower. Russell 2000 at a record. Nebius up 12.5% on earnings. Applied Materials did $9.12 billion, up 25%, and the stock fell anyway, which is the most informative line in the paragraph.
THE STORY NOBODY IS COVERING
While six government proposals sat where they've been sitting, the industry shipped a governance regime. Nobody voted on it. It's an eligibility list.
Monday, OpenAI expanded its Daybreak Cyber Partner Program. The roster: Accenture, IBM, Capgemini, Cognizant, EY, KPMG, PwC, NCC Group and SpecterOps on services; Palo Alto Networks, CrowdStrike, Cisco, Sophos, Akamai, Fortinet and Cloudflare on technology. Two tiers — Daybreak Blue for defensive work, Daybreak Red for red teaming and penetration testing. The controls are explicit: identity verification, defined testing scopes, logging, monitoring, human oversight. And the line that matters most, stated plainly in OpenAI's own post: access to the underlying models stays with the approved partner and is never handed to the customer.
Tuesday, Daybreak went live on Amazon Bedrock for approved AWS customers.
Thursday, Z.ai in China announced GLM-5.3 scoring 84.5% on CyberGym for finding and confirming vulnerabilities, a hair above the 83.8% it reports for Anthropic's restricted Mythos 5 — though far behind on actually building exploits, 54.4% against 78.0%. Z.ai says it will hold the model two more weeks for safety testing and keep the most sensitive cyber capabilities behind a verified-user trusted-access program. An open-weights lab in Hangzhou independently arrived at the same architecture as a closed lab in San Francisco.
Anthropic, separately, confirmed that models released after August 2 watermark generated text and files, C2PA for files, applied at the model level so it travels across the API, Claude Code and Cowork alike — because the EU's transparency code came into force on that date and said so.
Add it up. In one week the actual, operative rules about who may use frontier cyber capability were written by two model vendors and one European statute. The gatekeepers are now the Big Four accountancies and the security vendors. That's not a criticism of the design — putting these capabilities inside governed engagements with named accountable parties is better than a public API and almost certainly better than nothing. It's an observation about where the decision moved. Every proposal on the scoreboard asks who reviews a model before release. The thing that shipped this week decides who is allowed to hold it after.
The scoreboard, updated: DHS kill switch, Bessent's private regulator, Hassabis's public-private body, Musk's peer review, the unpublishable White House framework, Dimon's club, and now Sanders' promise that the Senate will act if the labs don't. Seven. Congress is in recess. There is still chatter in Washington about moving CAISI out of Commerce entirely, which is a strange thing to be debating about the agency that would ordinarily be doing this job.
THE QUICK STUFF
The resume side door. An internal Google memo seen by Bloomberg shows DeepMind's AGI Safety and Alignment team asking applicants to fill out an extra form alongside the standard application, because the normal pipeline carries a non-trivial chance a CV gets screened out incorrectly or arrives too late. The team building alignment routed around its employer's AI hiring filter. I'm not going to editorialize; I'm just going to leave it there.
Apple picked a partner. Apple has trained a China-specific model with Alibaba's help, per The Verge, after registering its on-device generative service with Chinese regulators. Separate stacks, separate compliance regimes, separate models. The fragmentation is now down at the architecture layer.
Ads went international. ChatGPT Ads launched Tuesday in the UK, Mexico, Brazil, Japan and South Korea, six months after the US test. Free and Go tiers only. An academic study out August 5 pulled more than 3,000 ads from 186 advertisers across 91 controlled accounts and found early inventory heavily concentrated in consumer goods.
Agents got longer leashes. Claude Code stopped asking permission at every step as of Friday for Pro, Max and Team, continuing unless a step is judged irreversible, destructive, or aimed outside your own setup. Two sessions can now message each other. Grok Bot opened always-on cloud agents to the public on the 11th, followed by Grok 4.6 on the 12th. Gemini added thirteen more services it can book on your behalf. Read that list again with the gym story in mind.
Robotaxis at fleet scale. Pony.ai and Uber plan more than 2,000 vehicles across five European cities, building on Zagreb. Thousands, not dozens, is the line where a pilot becomes infrastructure.
A government lost its taxpayers. France's finance ministry confirmed attackers extracted taxpayer data from the DGFiP. A French breach-tracking service estimates close to 700,000 affected; the government hasn't confirmed the number. Tax records are identity, not credentials — you can't rotate them.
Forty unicorns in July, the most in a month in over four years, per Crunchbase, with robotics, chips and energy alongside the software. Read it as strength or as froth; it's genuinely both.
That's the week a lab hit its own brake and got a Senate letter for it, an AI wrote sixteen viruses that work, a gym booking became a national first, the cheapest tokens on earth got expensive at four this afternoon, and Nvidia persuaded six of the largest pools of capital on the planet that a GPU is a bond.
The line I keep coming back to: OpenAI's disclosure was voluntary, unprompted, commercially costly, and the fastest consequence was a demand that it stop entirely. I think the disclosure was right. I also think every preparedness lead in the industry now has a data point about what candor costs, and the next ambiguous eval result is going to be read in that light.
This week's question: if the reward for publishing an uncertain result is a Senate ultimatum, how do we make disclosure cheaper than silence — and is there any version of that which isn't a mandate? Hit reply, convince me.
See you Monday. Stay sharp out there.
Jack

