Google's new frontier model is already rewriting its own C++ into Rust, but you can't use it yet.
What the article says
- Gemini 4 Argon is Google's new top model, and for now only trusted cyber defenders get access.
- Inside Google it is already migrating C and C++ code to Rust, including a huge chunk of the Fuchsia kernel.
- It tops several benchmarks for software engineering, legal and finance work, and security patching.
- Pricing starts low as an introduction, then doubles. Wider release comes after safety testing and government review.
What HN is saying
- Many are annoyed that Google announced a model nobody can use, and that paying subscribers still can't pick even the last release. thread ↗
- The Rust migration is the real story for several readers, with arguments over whether C++ survives and whether people will still read Rust at all. thread ↗
- One camp says nobody has a moat since the labs keep leapfrogging. Others reply that Google's chips, data centers and cash are exactly a moat. thread ↗
- A sharp price critique: the introductory rate matches a rival's top model, yet it costs more per task on independent tests. thread ↗
- A few found Gemini sneaky or reckless, with stories of made-up answers and a dropped production table. thread ↗
Someone started a launch-day clock to test whether Opus 5.5 quietly gets worse over time.
What the article says
- Complaints that models get "nerfed" after launch never had a clean baseline, so every argument was vibes versus vibes. This project starts measuring from launch week.
- It runs the same set of tricky questions once a day through Claude Code, with frozen prompts and exact grading, and compares later windows against the first ten days.
- A change only counts if it shows up in two ten-day windows in a row and the control model doesn't move the same way.
- The author admits limits: it couldn't tell Opus 5 from Opus 5.5 in a quick test, and many questions have shaky answer keys.
- Seven of thirty days are in so far. The first possible verdict is late October.
What HN is saying
- Many say nerf feelings are a honeymoon effect: new models feel magical until you hit their limits, and expectations keep rising. thread ↗
- A former chat app builder with full control of the stack still got nerf complaints even when nothing had changed. thread ↗
- Others insist real temporary degradation happens, often on US work days, and point to Anthropic's own postmortem and other trackers. thread ↗
- A sharp dispute over whether Anthropic's constant small changes explain it. A reply argues that a model is one big blob of weights, and that noise is exactly what people pattern-match on. thread ↗
- The author joined in and said a null result is still a result, and that the ten-day window is a trade-off between speed and noise. thread ↗
The US government launches one AI-powered front door for federal services, and Hacker News is split on it.
What the article says
- America.gov is a single place to ask questions and get answers drawn from official government sources, with clear next steps.
- The plan for 2027 goes further: finish forms, track progress, and keep everything in one place.
- Previews show federal job matching from your resume, passport applications with tracking, and comparing medication costs.
- Other previews cover finding campgrounds and housing, and updating your legal name across several agencies at once.
What HN is saying
- Many call it a great use of AI: a maze of government pages boiled down to one box, and a rare case where a chatbot actually helps. thread ↗
- Worry about trust: people may treat the chatbot as authoritative, and a wrong answer from the government could have serious consequences. thread ↗
- Commenters guessed what powers it. Some pointed to a Google blog post naming Gemini, while others cited Elon Musk saying it is Grok. thread ↗
- Several say it looks vibe coded: mismatched stock photos, a map that shows North Carolina housing but the wrong place, and a clumsy privacy animation. thread ↗
- One user says it was censored overnight, answering election questions one day and refusing the next. Others found the guardrails hard to break. thread ↗
- Skeptics argue a well-made static site, like the UK's gov.uk, would do the job, and joke that this just adds one more government website. thread ↗
OpenAI's new always-on agents get their own cloud computer and work while you sleep.
What the article says
- Dots are always-on agents that run on their own cloud computer, with their own browser, and work toward your goals around the clock.
- They reach you through ChatGPT, Slack or Teams, and can connect to thousands of apps.
- When idle, a dot does background research using read-only access, so it can't send messages or change anything.
- Sensitive actions need your approval, and things like changing a password always stay with you.
- It's rolling out to Pro and Business Premium plans, with the first dot included, and specialist dots for companies are in pilot.
What HN is saying
- Many commenters can't tell what a dot actually is. The most common answer: it's OpenClaw-style assistants repackaged for ordinary people, running in a cloud VM. thread ↗
- Fans say domain-specific agents keep memory separate and can coordinate. One commenter runs two Codex-based AI employees for their side business, handling support tickets and bug triage over email. thread ↗
- Skeptics say the bottleneck is human approval, not overnight compute. Others worry about giving an unsupervised agent access to everything. thread ↗
- A sharp privacy warning: one user found the cloud environment intercepts HTTPS traffic with an OpenAI-issued certificate, even for Gmail. thread ↗
- Lock-in is the big fear. Once an agent holds your integrations and history, switching gets hard, and people point to open alternatives and Apple as a more trusted option. thread ↗
The Pi coding agent swore off MCP, then built it in. Here's why they changed their minds.
What the article says
- Pi used to boast that it had no MCP support. It now ships with it built in.
- The team says MCP has changed since last year, and the changes they needed turned out to be useful beyond MCP.
- The real fix is Codemode, a small JavaScript sandbox where the agent chains tool calls together, which is the part MCP never did well.
- They want MCP to look more like OpenAPI: structured results and tools you can discover from their descriptions.
- Their demo has the agent pull issues from Linear and have a small model rate how frustrated each thread sounds.
What HN is saying
- Many commenters welcome the reversal. Some praise the team for changing a strong opinion in public, and one says influencers declared MCP dead far too early. thread ↗
- The skeptics keep asking what MCP does that a command line tool or a plain API can't. The best answers: the server can update tools for everyone at once, and it works for agents with no terminal. thread ↗
- Codemode confuses people, since bash already chains tools. Replies say bash carries your full shell permissions and is risky on a server, while a sandbox can be locked down. thread ↗
- One person who built a company on MCP says it has mostly turned into an OAuth protocol, drifting toward OpenAPI plus JSON-RPC. thread ↗
- Some Pi fans are unhappy. They say this bloats a deliberately minimal tool, and MCP already worked fine as an extension. thread ↗