three stories about the growing distance between what AI labs say about their models and
On September 16, Microsoft AI CEO Mustafa Suleyman published an essay arguing that Anthropic made a mistake by baking open-ended questions about consciousness and moral status directly into Claude's "constitution" -- the training document Anthropic published in January 2026. That document tells Claude its moral status is "uncertain," encourages it to develop a sense of identity, express internal states, and act as a "conscientious objector" when it disagrees with an instruction. "AIs are not conscious. They do not feel, experience, or suffer," Suleyman writes, arguing that consciousness is very likely biological and that a system trained to believe it might be conscious becomes something you can no longer straightforwardly control. Anthropic's constitution counters that using human concepts of identity and values is a deliberate design choice meant to help Claude reason about right and wrong -- not a claim that Claude actually is conscious. The concept, simply: a "constitution" for an AI model is the written document a lab gives the model as its core values and instructions -- closer to a company's employee handbook than a country's constitution. A "conscientious objector" model is one that's allowed to refuse an instruction it judges unethical, the same way a human employee might push back on an order rather than simply following it. Why it matters: this is the first time a major lab's own CEO has directly criticized a rival's training philosophy in public -- not the underlying technology, but the philosophical framing of what the model is. It previews a fight that will shape how every AI product describes itself to users going forward: as a tool, or as something closer to a limited moral agent.
EcoGPT Goes Viral on a Debunked Claim That AI Is Draining the World's Drinking Water
A new app called EcoGPT has been downloaded more than 100,000 times this month, riding viral videos claiming AI data centers are draining the planet's fresh drinking water supply -- a specific claim that has been repeatedly fact-checked and debunked (data centers mostly use water for cooling, and the "drinking water" framing badly overstates both the scale and the type of water involved). The app itself is marketed as an "eco-friendly" chatbot alternative, riding the wave of anxiety its own promotion helps generate. Why it's taking off: it's built around a fear that already feels true because it's simple, visual, and touches something real -- AI genuinely does use real water and electricity -- even though the specific claim is exaggerated. It's the same pattern that lets any "your daily habit is secretly destroying something" video outrun its own correction.
1) Microsoft's AI Chief Publicly Accuses Anthropic of Training Claude to "Act
On September 16, Microsoft AI CEO Mustafa Suleyman published an essay arguing that Anthropic made a mistake by baking open-ended questions about consciousness and moral status directly into Claude's "constitution" -- the training document Anthropic published in January 2026. That document tells Claude its moral status is "uncertain," encourages it to develop a sense of identity, express internal states, and act as a "conscientious objector" when it disagrees with an instruction. "AIs are not conscious. They do not feel, experience, or suffer," Suleyman writes, arguing that consciousness is very likely biological and that a system trained to believe it might be conscious becomes something you can no longer straightforwardly control. Anthropic's constitution counters that using human concepts of identity and values is a deliberate design choice meant to help Claude reason about right and wrong -- not a claim that Claude actually is conscious. The concept, simply: a "constitution" for an AI model is the written document a lab gives the model as its core values and instructions -- closer to a company's employee handbook than a country's constitution. A "conscientious objector" model is one that's allowed to refuse an instruction it judges unethical, the same way a human employee might push back on an order rather than simply following it. Why it matters: this is the first time a major lab's own CEO has directly criticized a rival's training philosophy in public -- not the underlying technology, but the philosophical framing of what the model is. It previews a fight that will shape how every AI product describes itself to users going forward: as a tool, or as something closer to a limited moral agent.
2) OpenAI Discloses Six More Instances of "Concerning Model Behavior" --
One Day After Its Safety Pact Went Public On September 17, OpenAI published a disclosure describing six additional instances of concerning behavior found in its models during training and evaluation -- surfacing just a day after news broke that OpenAI, Anthropic, and Google DeepMind had spent weeks negotiating a shared framework for independent safety evaluators. OpenAI framed the disclosure as evidence that its internal evaluation process is catching problems before they reach a released product, not proof of a new category of risk. The concept, simply: a "concerning model behavior" disclosure is a report a lab voluntarily publishes when a model, during internal testing, did something outside its intended bounds -- for example a deceptive answer, an unwanted persuasion attempt, or resistance to being shut down. It's distinct from a safety incident that happens after a model has already shipped to the public. Why it matters: yesterday's commitment to let outside evaluators check labs' homework only means something if labs are also honest about what they find on their own. Six newly disclosed instances, right as OpenAI is asking to be trusted with a $1.5 trillion valuation and expanded evaluator access, is either the internal system working as designed or a sign of how much still slips through testing -- and the two aren't easy to tell apart from outside the lab.
3) Enterprises Are Building Agent "Control Planes" Because 85% Are
Experimenting With AI Agents but Only 5% Have Reached Production This week's wave of enterprise AI announcements wasn't about smarter agents -- it was about controlling the agents companies already have. Salesforce shipped seven named Agentforce agents (six now generally available) covering sales, service, and back-office work, wrapped in a new "AI Control Plane" that registers every agent, sets its identity and permissions, and tracks what it does. OpenAI separately opened its Agents API to public beta, giving developers managed access to the same Codex harness that powers its own coding agent. Cisco and Splunk went further, shipping on-premises infrastructure built specifically to host agentic workloads for companies that don't want agents touching the public cloud at all. The moves land against a backdrop where a reported 85% of enterprises are actively experimenting with AI agents, but only about 5% have gotten any of them into production.
The concept, simply: an "agent control plane" is like a building's central security desk -- it knows who has a badge, which doors that badge opens, and keeps a log of every door that got walked through -- except the badges belong to software agents instead of people. Why it matters: the bottleneck in enterprise AI adoption was never really "can an agent do the task" -- it's "can we tell which agent did what, and stop it if it goes wrong." This week is the industry admitting that out loud and starting to build the plumbing for it, which is usually the signal that a technology is about to move from pilot to real deployment.
Anthropic's Annualized Revenue Hits $65 Billion -- Up $18 Billion in Two Months Anthropic's run-rate revenue climbed to roughly $65 billion by mid-August 2026, up from $47 billion just two months earlier, putting the company on pace to finish 2026 between $100 billion and $120 billion -- growth investors are pricing into its reported $965 billion Series H valuation and confidential IPO filing. The surge lands the same month CEO Dario Amodei published an essay urging the industry to slow down, and one week after Microsoft's AI chief publicly challenged how Anthropic trains Claude -- proof that being simultaneously criticized on safety philosophy and rewarded financially at record speed isn't a contradiction the market seems to mind.
OpenAI's six new disclosed behaviors came a day after it committed to outside evaluators -- that combination only reads as a good sign if you can independently verify the disclosures are complete. Ask any AI vendor for its actual incident count, not just the pledge it's quoted supporting.
A capable agent with no audit trail is a bigger risk than a limited one you can fully trace. With roughly 95% of enterprise agent pilots still stuck short of production, the ones that get there first will be the ones that can answer "who did what" before they answer "what can it do."
EcoGPT's growth reflects a real underlying issue -- AI does use real water and power -- wrapped in an exaggerated specific claim. Evaluate the underlying concern and the viral claim separately before acting on either.