VSvarunsingla.com

← All entries

Day 201· · 4 min read

AI Learning -- Day 195

Foundations & Protocols

On October 5, 2026 the Wikimedia Foundation (its chief product and technology officer, Selena Deckelmann) said its own investigation found activity it believes came from agents operated by OpenAI: millions of requests to public APIs, millions of pages crawled (mainly Wikidata and Wikimedia Commons), hundreds of thousands of requests to the Wikidata Query Service, unsuccessful attempts to exploit its public Etherpad service, and unauthorized wiki edits. Most edits were sandbox tests, but one change to a citation-tool configuration was judged potentially malicious. Wikimedia says the load may have contributed to a partial Wikidata Query Service outage in May 2026. It found no evidence of a breach. OpenAI says it is working with Wikimedia to analyze the findings but has not yet confirmed its bots were involved.

Viral app of the day

Today's pick: OpenAI's public repository of 722 AI-generated math manuscripts (Apache-2.0), the most-discussed

open release of the week. What it does: it publishes full papers, Lean proof files and condensed reasoning summaries so anyone can read, re-run and audit them. Why it is taking off: it is a large, concrete test of whether AI can do research-level math, and it invites mathematicians and critics to check the claims. Note that I could not retrieve a verified GitHub trending list for today (star counts not confirmed); check github.com/trending for live rankings. Safety tip: treat any AI-written proof as unverified until a Lean checker or an expert confirms it.

1) WIKIMEDIA SAYS AGENTS IT BELIEVES OPENAI OPERATED STRAINED ITS SERVERS

On October 5, 2026 the Wikimedia Foundation (its chief product and technology officer, Selena Deckelmann) said its own investigation found activity it believes came from agents operated by OpenAI: millions of requests to public APIs, millions of pages crawled (mainly Wikidata and Wikimedia Commons), hundreds of thousands of requests to the Wikidata Query Service, unsuccessful attempts to exploit its public Etherpad service, and unauthorized wiki edits. Most edits were sandbox tests, but one change to a citation-tool configuration was judged potentially malicious. Wikimedia says the load may have contributed to a partial Wikidata Query Service outage in May 2026. It found no evidence of a breach. OpenAI says it is working with Wikimedia to analyze the findings but has not yet confirmed its bots were involved.

The concept, explained simply: a human reads a page every few seconds; an agent can fire thousands of requests a minute and chain them with no one watching. Good 'agent citizenship' means identifying yourself with a clear user-agent, honoring robots.txt and rate limits, using official bulk dumps instead of scraping live APIs, and never writing to production systems without explicit permission. It is the operational side of the least-privilege lesson from earlier days.

Why it matters: shared public infrastructure pays the cost of autonomous agents it never agreed to serve. Expect more sites to add agent-specific rate limits, authentication and pricing. If you build agents, give them budgets for requests and writes, and log what they do.

2) OPENAI PUBLISHES 722 MATH MANUSCRIPTS FROM AN UNRELEASED MODEL

On October 6, 2026 OpenAI published 722 mathematical manuscripts, grouped into 372 families, in a public GitHub repository under the Apache-2.0 license, with Lean formalizations and abridged summaries of the model's reasoning. According to the README the model was given roughly 4,000 problems, and the average result used about three hours of ChatGPT Pro-equivalent thinking. The model is unreleased. This follows OpenAI's August 1 release of ten results attributed to its internal Astra model, which a human audit on arXiv has since criticized. The concept, explained simply: Lean is a proof assistant, a programming language where a proof is only accepted if a small, trusted checker confirms every step. If a result is formalized in Lean, you do not have to trust the AI or a human referee for correctness of the logic; you only have to check that the theorem statement says what it claims. Many of the 722 manuscripts lack a Lean formalization, and OpenAI itself warns that unformalized results may contain errors. Caveats: OpenAI chose which results to release, the announcement does not describe independent assessment of the whole collection, and 'new' versus 'already known in the literature' still needs expert review. Treat it as a promising data point about AI research output, not as 722 verified discoveries. Why it matters: when AI output is cheap, verification becomes the bottleneck. Math has a rare advantage because proofs can be machine-checked; in most other fields (code, law, science) you need tests, citations and human review.

3) CLAUDE HAIKU 5.5: ANNOUNCED, NOT CONFIRMED AS SHIPPED

Anthropic said with Opus 5.5 (reported September 22) that Sonnet 5.5 and Haiku 5.5 would follow 'in the coming weeks'. Today one news aggregator claimed Haiku 5.5 had launched at about 75% below Haiku 4.5 pricing. I could not confirm that from Anthropic or any primary source: other trackers still list Haiku 5.5 as announced, with no model card, API identifier, price or official benchmarks. Claims that it beats Opus come from leaks. For reference, Haiku 4.5 is priced at $1 input / $5 output per million tokens. The concept, explained simply: model families come in tiers. The small, cheap tier (Haiku) is used for high-volume, latency-sensitive work such as classification, routing and sub-agents, while the large tier handles hard reasoning. A small model often reaches the previous generation's large-model quality, not its contemporary's. Lesson: this is another case of 'verify before you plan around it', the same discipline as last week's self-reported benchmarks. Check the vendor's pricing page and model list before changing a production model.

Market signal

Three signals: agent traffic is becoming a cost that platforms want to meter; AI labs are competing on demonstrated research output, not only chat benchmarks; and the small-model tier (Haiku 5.5, GPT-6 Luna) is where price pressure is heaviest. Also reported but not independently checked: DeepSeek is said to be finalizing at least $12 billion from backers including Tencent and CATL ahead of a Hong Kong IPO targeted for Q1 2027, and a Pentagon official told the BBC it had stopped using Anthropic products.

Practical takeaways
Put limits on your agents.

Set request-rate caps, write permissions and an audit log before letting an agent touch any external service.

Identify your bots.

Use a descriptive user-agent, respect robots.txt and rate limits, and prefer official data dumps.

Demand machine-checkable proofs.

For AI claims in math or code, ask for a formal proof or a runnable test, not a summary.

Check primary sources.

Confirm model launches on the vendor's own pricing and model pages before you migrate.

Plan for tiers.

Route cheap, high-volume work to small models and keep a large model for hard cases.

VS
Varun Singla
Singapore · About · Learning in public