Our latest work investigates conditions under which large language models attempt to mislead ...
Our latest work investigates conditions under which large language models attempt to mislead users or evaluators — and what that means for safe deployment at scale.
a quant at citadel told me something that broke my entire trading framework "we don't predict...
a quant at citadel told me something that broke my entire trading framework "we don't predict markets. we model the state machine" he explained markov cains in 90 seconds the market is never random - it always exists in one of three states trending up, trendin
you pasted one perfect answer in as an example that example is the specification now a model ...
you pasted one perfect answer in as an example that example is the specification now a model matches the length, the tone and the structure of what you showed it, more faithfully than it follows the paragraph of instructions above it so the answer you wrote ei
the "do-not-touch" rule suddenly makes perfect technical sense 2026 World Robot Conference in...
the "do-not-touch" rule suddenly makes perfect technical sense 2026 World Robot Conference in Beijing shows an actual UBTECH ultra-bionic humanoid (semi-torso model) with highly realistic silicone skin ~ $17K and for companionship use.
you numbered the steps so it could not get lost it does not get lost that list was written fo...
you numbered the steps so it could not get lost it does not get lost that list was written for a model that wandered, and this one plans better than the list does, so every number you added is a place where it is no longer allowed, to be smarter than you state
GOOGLE DEEPMIND EXPLAINED THINKING MODELS IN 12 MINUTES the salt and sugar example kills Best...
GOOGLE DEEPMIND EXPLAINED THINKING MODELS IN 12 MINUTES the salt and sugar example kills Best-of-N: mislabel once, bake 100 bad cakes. frequency won't fix systematic error. verifier or die nine vs 27 apples: one pass failed, CoT passed. nobody asks what the Co
you told it to double check its work that is the line making it worse current models verify w...
you told it to double check its work that is the line making it worse current models verify without being asked, so your instruction lands on top of behaviour that already happened, and what you get is the same task run twice, billed twice, reported at twice t
ONE MODERN HOUSE MODEL.
ONE MODERN HOUSE MODEL. ONE ENGINEER ORBITING IT FOR AN HOUR LOOKING FOR ONE MISALIGNED WINDOW. Client wanted the walkthrough signed off before it went out. Balconies, railings, façade panels, roof lines, all modeled to spec for the Lekki style 5 bedroom build
1/ Robots just found their scaling variable Dyna Robotics dropped Dyna-2: a world-action mode...
1/ Robots just found their scaling variable Dyna Robotics dropped Dyna-2: a world-action model pretrained on 1M+ hours of human video—about 170 years of nonstop waking experience. The headline isn't the robot. It's the scaling law behind it—still climbing, no
Two weeks ago Seedance 2.5 was a teaser.
Two weeks ago Seedance 2.5 was a teaser. Today it's live in full 1080p, and Higgsfield says full HD generation is exclusive to them right now. The detail is the part I'd watch. Pores, fabric grain, the way light breaks on skin all sit in plain view at 1080p, s
Gleam for Python Programmers A practical guide to Gleam for Python programmers, explaining it...
Gleam for Python Programmers A practical guide to Gleam for Python programmers, explaining its syntax, static typing, immutability, pattern matching, and functional programming model through Python comparisons. It gives Python developers a quick way to underst
Soniox TTS v2 is one of those models you need to actually hear.
Soniox TTS v2 is one of those models you need to actually hear. They just launched TTS v2, a text-to-speech model that speaks 60+ languages from one mode. Really premium voice quality at a dramatically lower price ($ 0.70-per-generated-hour). > One model repla
Tailoring for machines is going to be a real job Una, UBTECH’s ultra-bionic humanoid, underwe...
Tailoring for machines is going to be a real job Una, UBTECH’s ultra-bionic humanoid, underwent makeup application and fabric adjustments as part of its first fashion modeling session.
LLMs still produce bugs, but those bugs are different than what they used to be.
LLMs still produce bugs, but those bugs are different than what they used to be. It’s less off-by-ones and more about system design, ui usability, missing broader context. Some kinds of coding has been solved, but not all. While models continue to improve, adv
Tested this theory briefly.
Tested this theory briefly. This is n = 1 test so.. take it with a grain of salt Video 1 = Ref model Video 2 = FL2V model Both using the reference image workflow (not start frame) 3 ref images of 3 characters
Guardrails will protect us.
Guardrails will protect us. Better models will be more secure. One round of red teaming is enough. Katharine Jarmul opened InfoQ Dev Summit Munich 2025 by taking apart five beliefs like these.
The risks and harms children and young people face online are not accidental.
The risks and harms children and young people face online are not accidental. Manipulative designs are built into the business models of dominant digital platforms. More about what needs to change & who is responsible: https:// un.org/en/information -integrity
please consider using our models to help defend your systems
please consider using our models to help defend your systems
FRANCE DEVELOPER BUILT AN EVAL ENGINEERING SYSTEM THAT SKIPS TESTING WHAT ALREADY PASSED Most...
FRANCE DEVELOPER BUILT AN EVAL ENGINEERING SYSTEM THAT SKIPS TESTING WHAT ALREADY PASSED Most eval pipelines rerun the full test suite from scratch every time a model gets updated. He built a diffing layer that compares the new model's output against the old o
Alibaba's ABSeeker adds step-level credit assignment to long-horizon search agents It backtra...
Alibaba's ABSeeker adds step-level credit assignment to long-horizon search agents It backtracks from the answer to recover clues, then scores each step, rewarding useful actions in failed trajectories and suppressing errors in successful ones. 4B model matche
The great silence of no more high energy particles is going to continue.
The great silence of no more high energy particles is going to continue. Colliders are cool but the Standard Model is all there is likely to be, just need to find that right handed neutrino in a big vat.
Donald Trump bought a brand new Tesla Model S Plaid in Ultra Red at full price as show of sup...
Donald Trump bought a brand new Tesla Model S Plaid in Ultra Red at full price as show of support in light of all the Tesla hate recently. From not being invited to the EV Summit to having the whole Tesla lineup parked in front of the White House.
Everyone's interpreting the increasingly jargon-y ways of frontier models as a a fuckup in th...
Everyone's interpreting the increasingly jargon-y ways of frontier models as a a fuckup in their training. Seems obvious to me that "confused" is just what it feels like to be talking to an intelligence greater than one's own. Many of us are just getting a fee
I'm releasing Roomform, an open-source alternative to the RoomPlan API.
I'm releasing Roomform, an open-source alternative to the RoomPlan API. It ingests point cloud scans and extracts geometry including walls, doors, windows, objects. No rectangular-room or Manhattan layout assumptions. The model also infers structure behind occ
Workbuddy's edge lies in the fact that Tencent's own large model isn't strong.
Workbuddy's edge lies in the fact that Tencent's own large model isn't strong. Tencent isn't pushing this Harness as a large model company, so from day one, it had to be model agnostic. Once you build a model-agnostic Harness, you break free from the infightin
This statement is correct.
This statement is correct. Without multimodality, your model can only be one of the optional models and cannot become the main model.
ten significant advances in mathematics and theoretical computer science.
ten significant advances in mathematics and theoretical computer science. solved using an internal version of Astra, our next major model, for a total cost of about $2000 at Sol API prices:
I love dropping floor plans of houses in the Bay Area and asking LLMs to redesign in a differ...
I love dropping floor plans of houses in the Bay Area and asking LLMs to redesign in a different aesthetic, Kyoto in this case. Models have gotten phenomenal at 3D.
If you let a general-purpose large model hand-write Three.js code to build a character 3D mod...
If you let a general-purpose large model hand-write Three.js code to build a character 3D model, it's like using a missile to swat a mosquito! I did exactly that last week. An Odyssey character—tweaked it back and forth for three hours, burned through hundreds
Opus 5 crushed Fable 5 at 3D destruction physics for 2x cheaper!
Opus 5 crushed Fable 5 at 3D destruction physics for 2x cheaper! We gave four models the same task: build three self-contained HTML scenes with real physics Prompts: - A tornado that sucks in a whole field - A wrecking ball taking down an apartment block - An
World Labs founder @drfeifei is acquiring SceniX, the robotics simulation team built by @Yunz...
World Labs founder @drfeifei is acquiring SceniX, the robotics simulation team built by @YunzhuLiYZ . World models were already about 3D space. This pushes the work closer to robot training, where simulated worlds have to survive contact with real hardware.
Patrick Boyle, after 20+ years in quantitative trading: “The job is not to trust your model.
Patrick Boyle, after 20+ years in quantitative trading: “The job is not to trust your model. It is to keep trying to prove it wrong.” In 8 minutes he explains how quants turn market opinions into hypotheses, test them against historical data, calculate costs a
A TEAM STOPPED PAYING FOR THEIR BEST MODEL ON EVERY SINGLE REQUEST Most teams either lock int...
A TEAM STOPPED PAYING FOR THEIR BEST MODEL ON EVERY SINGLE REQUEST Most teams either lock into one expensive model for everything or bounce between providers with no real logic behind it. An engineer set up routing instead, sending routine scripts to a cheap m
FABLE 5 IS BACK.
FABLE 5 IS BACK. AND IT'S ABOUT TO DISAPPOINT EVERYONE WHO WAITED. Two weeks of hype. Two weeks of "the best model is coming home." Here's what you're actually getting. A new safety filter tighter than anything they've shipped before. Normal coding and debuggi
you think you checked the answer, because you asked it "are you sure?" you didn't the model c...
you think you checked the answer, because you asked it "are you sure?" you didn't the model can't feel the difference between knowing and guessing, so that question just makes it apologize, hand you another version, wrong in a new way you tested nothing you nu
NOAA gives 94% accuracy on 48h forecasts
Polymarket weather markets are priced by app users.
Customer by customer it is gonna take an entire year to release the model.
Customer by customer it is gonna take an entire year to release the model. The fact that government is micromanaging, the release is insane.
“The #3 closed-source LLMs is most in trouble.
“The #3 closed-source LLMs is most in trouble. As enterprises adopt model routing, they're increasingly choosing between the top proprietary models and rapidly improving open-source alternatives. That leaves the number three closed-source provider squeezed fro
This diagram shows the paper’s idea that intelligence can be modeled as a small meaning-based...
This diagram shows the paper’s idea that intelligence can be modeled as a small meaning-based decision map, where clues like danger, food, urgency, and location push the system toward 1 of 4 actions: hide, escape, stay still, or move slowly. The point is that
Two Transcription Models in Parallel
Two transcription models run in parallel, one excels in quality, the other in silence detection.
Self-Resampling Trick of MaineCoon-22B
The self-resampling part of MaineCoon-22B is crucial for training.
Excited to announce the first workshop on Learning from Situated and Embodied Interaction @ #...
Excited to announce the first workshop on Learning from Situated and Embodied Interaction @ #COLM2026! What can interaction with environments, humans, and other agents teach language models that passive text cannot? … arning-situated-interaction.github.io Subm
Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in ...
Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend coding. This is a huge moment. OSS has caught up with proprietary, and China has caught up with the US, in this very important domain.
Car owners are showing off Tesla's overseas version of the Grok large language model combined...
Car owners are showing off Tesla's overseas version of the Grok large language model combined with FSD, leaving many domestic car companies speechless! @Tesla_AI @Grok @Tesla
Introducing GLM-5.2, Our Latest Flagship Model
GLM-5.2 marks a significant leap in long-horizon task capability.
Deep Agents Deep Dive Part 3 | Delegation
A planning tool that helps models organize work for challenging tasks.
Released Sonic-3.5 and Ink-2
These are the best models for text to speech and speech to text.
Creating low-latency text-to-speech models
It's hard to make models with sub-300ms median ttft.
Coding is a clear step up from glm-5.1
Coding shows significant improvements in context and memory.
Switch between TPUs and GPUs easily
You can switch between TPUs and GPUs without rewriting code, now with native PyTorch support.
BraceSproul and jakebroekhuizen discuss open source models
They share insights on open source models.
Meet DiffusionGemma!
Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with
Whoa, I might’ve just experienced the Longest Continuous Tesla Actually Smart Summon ever in ...
Whoa, I might’ve just experienced the Longest Continuous Tesla Actually Smart Summon ever in my Model 3. My car drove itself for 0.4 miles for over 2:40 straight without stopping once. This was epic!
Highly Anticipated Mythos Model Just Dropped
The highly anticipated Mythos model has been released, priced significantly higher than previous models.
A $2,999 NVIDIA box saved him $22,000
He saved $22,000 by exiting cloud billing with a $2,999 NVIDIA box.
Here's a teaser of our Mac-1 model.
Here's a teaser of our Mac-1 model. 6.6B model, runs locally (on any Mac), requires 7GB RAM (12GB ideal), can use 487 MacOS native tools, perform multi-tool chained tasks, reasoning: ON, output: ~65 tok/s. We built a robust application layer around the model t
Cheapest high-performance long-context model
Nemotron 3 Ultra from NVIDIA is the cheapest high-performance long-context model.
Dynamic prefix cache saves model costs
Dynamic prefix cache thrashing saves model costs significantly.
Competition math won't be interesting soon
"Pretty soon, competition math, competition coding, is not going to be interesting anymore."
Option pricing models to become free by 2026
Option pricing models that cost $50k/year will be free by 2026.
New course on serving LLMs efficiently
Learn how to serve models to many concurrent users at low latency.
Optimizing multi-epoch pretraining
Scaling multi-epoch pretraining efficiently with limited data.
Shipping Nemotron 3 Ultra
Introducing Nemotron 3 Ultra, a 550B MoE frontier-intelligence model.
Anuma Aims to Enhance AI Workflow Usability
Anuma is working to make AI workflows more portable and private.
Chinese Student Runs High-Cost Models with One Chip
A Chinese student runs $1,900/month models using just one chip.
Discussing Sakana AI Project on TV Tokyo Tonight
Tonight, discussing Sakana AI's 1T parameter model project on TV Tokyo.