Mistral's biggest open-weight model yet, trained and hosted in Europe, and it won't refuse security work.
What the article says
- Mistral Large 4 is a huge multimodal model, open weights promised by the end of the month, with a public preview API available now.
- It was trained from scratch on Mistral's own European datacenters, and it will be served from a European deployment under European law.
- Mistral pitches it for cybersecurity. It tops a test of reproducing and patching real vulnerabilities, where some closed models score near zero because they refuse.
- On coding it ranks second in a blind human rating, behind only Claude Opus 5, and it beats several leading Chinese open models.
What HN is saying
- Many see a solid, useful model, especially for cybersecurity and as a European alternative to Chinese open models. thread ↗
- Others call it mediocre and months late, and read it as a grim sign for European AI. thread ↗
- Data point from a commenter at Plotly: on their analytics benchmark it is far cheaper than Mistral's last model and much more accurate, though a smaller Qwen may still beat it on price. thread ↗
- Sovereignty is the draw for some: trained and served in the EU matters to companies that won't hand product features to US suppliers. thread ↗
- A sharp question: if ~4,000 GPUs can get this close to the frontier, what are the giant US datacenters for? Replies say closed models are still ahead and compute also goes to inference. thread ↗
OpenAI dumps a pile of AI-written proofs of famous open math problems, and mathematicians are reeling.
What the article says
- OpenAI is publishing new results on open math problems from an internal model, all in a public GitHub repository.
- The release follows advice from an independent advisory group at the Institute for Advanced Study, with rules for revising and citing the papers.
- Many proofs come with Lean formalizations, so a computer can check them.
- OpenAI also shares summaries of the model's reasoning and says the average result took about three hours of ChatGPT Pro thinking.
- It plans to fund workshops on understanding the results and says it is working to release the model responsibly.
What HN is saying
- Several people with real stakes in the problems felt a mix of awe and grief. One commenter spent 24 years on Barnette's Conjecture, which now appears solved. thread ↗
- A sharp worry: this kind of progress only comes from tightly held, proprietary models at one or two labs, which nobody else can reproduce. thread ↗
- Skeptics say it is a win for formal verification, reinforcement learning and huge compute, not proof of broad insight across fields of math. thread ↗
- Fans picked out the surprises: the Unique Games Conjecture, faster integer multiplication, and a scheduling algorithm with an absurdly large exponent. thread ↗
- Doomers and non-doomers argued over what it means. The calmer view is that math is easy to verify and compute was enormous, so it may not carry over to the messy real world. thread ↗
JetBrains posted a loss while spending big. Is the IDE maker in trouble, or betting on AI?
What the article says
- JetBrains grew revenue modestly in 2025 but ended the year with a net financial loss. (Inferred from the title, description and comments, since the page text is thin.)
- Staff costs jumped by roughly a third, far faster than sales.
- Money spent on investing activities rose sharply, which suggests a big bet on something new.
- Total assets still grew, so this looks like heavy spending, not a company in distress.
What HN is saying
- Many commenters say the doom talk is overblown. Revenue is still growing and the extra spending looks like deliberate investment. thread ↗
- A common guess is that the money is going into AI products and cheap tokens to keep users from drifting to Claude and friends. thread ↗
- A lot of developers admit they now rarely read code in an IDE, so they have dropped JetBrains for lighter editors or the terminal. thread ↗
- Some blame management for chasing AI trends and killing products like Upsource, while others say JetBrains tools still have no real rival. thread ↗
- A wish list emerged for an IDE built for agents, with worktrees, parallel sessions and diff review. Editor caches also go stale while an agent edits. thread ↗
ChatGPT's fake New Yorker cartoons come with real cartoonists' signatures, and they are not happy.
What the article says
- A viral Dolly Parton cartoon carried the pen name of New Yorker cartoonist Brendan Loper. ChatGPT made it, not him.
- Reporters found more than 15 real cartoonists whose signatures the image generator reproduced, including Emily Flake and Saul Steinberg.
- Cartoonists call the signature their certificate of authenticity. Seeing it forged feels like impersonation, not just copying.
- OpenAI added a guardrail warning after being told, but ChatGPT still signs some cartoons with real names.
- Condé Nast says it never allowed anyone to train models on its cartoons.
What HN is saying
- Many say the cause is plain: training data links New Yorker style with a corner signature, so the model paints one on without understanding it. thread ↗
- Strong appetite for lawsuits. Commenters point to forged names as libel, trademark or right of publicity claims, especially since OpenAI charges subscriptions. thread ↗
- Pushback on the legal angle: drawing a fake signature for yourself is fine, and liability usually starts when someone publishes or sells it. thread ↗
- Skeptics see this as proof that AI companies copy first and ask later, while boosters' claims about reasoning take a hit. thread ↗
- Others note it is old news. Image models have always stamped fake marks, like Japanese seals on imitation woodblock prints. thread ↗
A new Western open-weight coding model is announced, but the weights are not out yet.
What the article says
- Reflection AI unveils Beam, its first open-weight model, built for coding, reasoning and agent work.
- It is a sparse model with 501 billion parameters, but only 23 billion are active at a time, which keeps running costs down.
- The team says it put huge compute into reinforcement learning, with no sign of the gains levelling off.
- It matches some larger open models on reasoning while using far less compute, but admits Kimi K3 is still ahead on raw capability.
- Weights and a technical report are promised later this month. For now there is only an early access sign-up.
What HN is saying
- Many say the benchmark chart hides stronger open models, and that DeepSeek V4.1 Flash looks smaller, cheaper and better on every measured score. thread ↗
- Defenders reply that the pitch is simply being Western, and that any new open-weight entrant is welcome even if it is not record-breaking. thread ↗
- Frustration that this is an announcement with no weights: show the files or stay quiet. One commenter calls it a marketing miss. thread ↗
- A sharp catch: the announcement calls the land-or-water puzzle brand new, but commenters say it has been around much longer. thread ↗
- Some wanted to know who Reflection is. The answer is a well-funded startup founded by ex-DeepMind researchers, unrelated to the 2024 Reflection 70B fiasco. thread ↗