If you let a general-purpose large model hand-write Three.js code to build a character 3D mod...
If you let a general-purpose large model hand-write Three.js code to build a character 3D model, it's like using a missile to swat a mosquito! I did exactly that last week. An Odyssey character—tweaked it back and forth for three hours, burned through hundreds
Opus 5 crushed Fable 5 at 3D destruction physics for 2x cheaper!
Opus 5 crushed Fable 5 at 3D destruction physics for 2x cheaper! We gave four models the same task: build three self-contained HTML scenes with real physics Prompts: - A tornado that sucks in a whole field - A wrecking ball taking down an apartment block - An
World Labs founder @drfeifei is acquiring SceniX, the robotics simulation team built by @Yunz...
World Labs founder @drfeifei is acquiring SceniX, the robotics simulation team built by @YunzhuLiYZ . World models were already about 3D space. This pushes the work closer to robot training, where simulated worlds have to survive contact with real hardware.
Patrick Boyle, after 20+ years in quantitative trading: “The job is not to trust your model.
Patrick Boyle, after 20+ years in quantitative trading: “The job is not to trust your model. It is to keep trying to prove it wrong.” In 8 minutes he explains how quants turn market opinions into hypotheses, test them against historical data, calculate costs a
A TEAM STOPPED PAYING FOR THEIR BEST MODEL ON EVERY SINGLE REQUEST Most teams either lock int...
A TEAM STOPPED PAYING FOR THEIR BEST MODEL ON EVERY SINGLE REQUEST Most teams either lock into one expensive model for everything or bounce between providers with no real logic behind it. An engineer set up routing instead, sending routine scripts to a cheap m
FABLE 5 IS BACK.
FABLE 5 IS BACK. AND IT'S ABOUT TO DISAPPOINT EVERYONE WHO WAITED. Two weeks of hype. Two weeks of "the best model is coming home." Here's what you're actually getting. A new safety filter tighter than anything they've shipped before. Normal coding and debuggi
you think you checked the answer, because you asked it "are you sure?" you didn't the model c...
you think you checked the answer, because you asked it "are you sure?" you didn't the model can't feel the difference between knowing and guessing, so that question just makes it apologize, hand you another version, wrong in a new way you tested nothing you nu
NOAA gives 94% accuracy on 48h forecasts
Polymarket weather markets are priced by app users.
Customer by customer it is gonna take an entire year to release the model.
Customer by customer it is gonna take an entire year to release the model. The fact that government is micromanaging, the release is insane.
“The #3 closed-source LLMs is most in trouble.
“The #3 closed-source LLMs is most in trouble. As enterprises adopt model routing, they're increasingly choosing between the top proprietary models and rapidly improving open-source alternatives. That leaves the number three closed-source provider squeezed fro
This diagram shows the paper’s idea that intelligence can be modeled as a small meaning-based...
This diagram shows the paper’s idea that intelligence can be modeled as a small meaning-based decision map, where clues like danger, food, urgency, and location push the system toward 1 of 4 actions: hide, escape, stay still, or move slowly. The point is that
Two Transcription Models in Parallel
Two transcription models run in parallel, one excels in quality, the other in silence detection.
Self-Resampling Trick of MaineCoon-22B
The self-resampling part of MaineCoon-22B is crucial for training.
Excited to announce the first workshop on Learning from Situated and Embodied Interaction @ #...
Excited to announce the first workshop on Learning from Situated and Embodied Interaction @ #COLM2026! What can interaction with environments, humans, and other agents teach language models that passive text cannot? … arning-situated-interaction.github.io Subm
Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in ...
Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend coding. This is a huge moment. OSS has caught up with proprietary, and China has caught up with the US, in this very important domain.
Car owners are showing off Tesla's overseas version of the Grok large language model combined...
Car owners are showing off Tesla's overseas version of the Grok large language model combined with FSD, leaving many domestic car companies speechless! @Tesla_AI @Grok @Tesla
Introducing GLM-5.2, Our Latest Flagship Model
GLM-5.2 marks a significant leap in long-horizon task capability.
Deep Agents Deep Dive Part 3 | Delegation
A planning tool that helps models organize work for challenging tasks.
Released Sonic-3.5 and Ink-2
These are the best models for text to speech and speech to text.
Creating low-latency text-to-speech models
It's hard to make models with sub-300ms median ttft.
Coding is a clear step up from glm-5.1
Coding shows significant improvements in context and memory.
Switch between TPUs and GPUs easily
You can switch between TPUs and GPUs without rewriting code, now with native PyTorch support.
BraceSproul and jakebroekhuizen discuss open source models
They share insights on open source models.
Meet DiffusionGemma!
Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with
Whoa, I might’ve just experienced the Longest Continuous Tesla Actually Smart Summon ever in ...
Whoa, I might’ve just experienced the Longest Continuous Tesla Actually Smart Summon ever in my Model 3. My car drove itself for 0.4 miles for over 2:40 straight without stopping once. This was epic!
Highly Anticipated Mythos Model Just Dropped
The highly anticipated Mythos model has been released, priced significantly higher than previous models.
A $2,999 NVIDIA box saved him $22,000
He saved $22,000 by exiting cloud billing with a $2,999 NVIDIA box.
Here's a teaser of our Mac-1 model.
Here's a teaser of our Mac-1 model. 6.6B model, runs locally (on any Mac), requires 7GB RAM (12GB ideal), can use 487 MacOS native tools, perform multi-tool chained tasks, reasoning: ON, output: ~65 tok/s. We built a robust application layer around the model t
Cheapest high-performance long-context model
Nemotron 3 Ultra from NVIDIA is the cheapest high-performance long-context model.
Dynamic prefix cache saves model costs
Dynamic prefix cache thrashing saves model costs significantly.
Competition math won't be interesting soon
"Pretty soon, competition math, competition coding, is not going to be interesting anymore."
Option pricing models to become free by 2026
Option pricing models that cost $50k/year will be free by 2026.
New course on serving LLMs efficiently
Learn how to serve models to many concurrent users at low latency.
Optimizing multi-epoch pretraining
Scaling multi-epoch pretraining efficiently with limited data.
Shipping Nemotron 3 Ultra
Introducing Nemotron 3 Ultra, a 550B MoE frontier-intelligence model.
Anuma Aims to Enhance AI Workflow Usability
Anuma is working to make AI workflows more portable and private.
Chinese Student Runs High-Cost Models with One Chip
A Chinese student runs $1,900/month models using just one chip.
Discussing Sakana AI Project on TV Tokyo Tonight
Tonight, discussing Sakana AI's 1T parameter model project on TV Tokyo.