FOUR MODELS.
FOUR MODELS. THREE TASKS. ZERO FIXES. everything looked the same at the start. but after one prompt, the difference became obvious. GPT-5.6, Claude Opus 5, Qwen 3.8 Max and Grok 4.6 had to build a premium product page, a playable browser game, and an animated
I ran 10 prompts blind through GPT-4o, Claude 5, and Grok 4.6.
I ran 10 prompts blind through GPT-4o, Claude 5, and Grok 4.6. Same task. No labels. Just the output. Results: > 3 tasks: GPT-4o won > 4 tasks: Claude 5 won > 3 tasks: Grok 4.6 won 2-point spread on benchmarks. My test confirms it. The model stopped being the
Anthropic Engineer Andrej Karpathy dropped a full 6-hour guide on "Intro and Deep Dive to LLM...
Anthropic Engineer Andrej Karpathy dropped a full 6-hour guide on "Intro and Deep Dive to LLMs and Tokenizer: • 00:00 - Intro to LLMs • 00:59:28 - Deep dive into LLMs (Karpathy method) • 03:31:43 - Let s build the GPT Tokenizer from scratch This Karpathy cours
GPT-5.6 Sol 1M in Codex.
GPT-5.6 Sol 1M in Codex. This used to only work for API keys, but we just flipped the switch and works for usage through ChatGPT accounts now too. The same warning applies, there is a reason the current context length is the default, we have tuned it to ~perfe
Here is how to enable a 1M-token context window in Codex for GPT-5.6 Sol.
Here is how to enable a 1M-token context window in Codex for GPT-5.6 Sol. Even though we have tuned the context limit in Codex to be set optimally when it comes to performance and cost, this is a common ask, so here it is documented. A larger context window le
deepseek 马上发布code版本,可以直接调用工具跟执行 马上国内可以调用电脑的版本就出来了 按照deepseek的便宜价格,应该会抢很大一部分GPT 的市场 目前公众号已经注册了...
deepseek 马上发布code版本,可以直接调用工具跟执行 马上国内可以调用电脑的版本就出来了 按照deepseek的便宜价格,应该会抢很大一部分GPT 的市场 目前公众号已经注册了,马上开始招测试用户
.
. @Toyota built their R&D GPT on Deep Agents to help employees search years of paint, seat, and manufacturing research with a single command. Here's their story.
do you understand what OpenAI just did?
do you understand what OpenAI just did? built a hacking-grade AI, then locked it behind an approval process for “trusted defenders” only it’s called Daybreak, and the headline model is GPT-5.6-Cyber. the numbers tell the story. on advanced cybersecurity tasks,
Developers weigh in on OpenAI's GPT-5.6 Sol, from database audits to Erdős problem solving, a...
Developers weigh in on OpenAI's GPT-5.6 Sol, from database audits to Erdős problem solving, as it takes on Anthropic's Claude Opus 5.
based on the error messages, we are about to get a new GPT model.
based on the error messages, we are about to get a new GPT model.
Fable 5 is great at under-defined feature building GPT 5.6 Sol is great at most workhorse tas...
Fable 5 is great at under-defined feature building GPT 5.6 Sol is great at most workhorse tasks Grok 4.5 is great where token efficiency matters, like PR review Opus 5 GPT 5.6 Luna Max is 80% of Sol at way less cost
翻到一条六年前发的、体验 GPT-3 后的推文,没想到这一晃就六年了,这下真的开始慌了。
翻到一条六年前发的、体验 GPT-3 后的推文,没想到这一晃就六年了,这下真的开始慌了。
I just realized I will be in San Francisco again on Sep 30th and that this is the exact date ...
I just realized I will be in San Francisco again on Sep 30th and that this is the exact date I predicted I would turn more bullish on OpenAI than Anthropic maybe they delay GPT-6 until DevDay on Sep 29th, which could technically be Sep 30th in Germany (if that
GPT-4 finished training four years ago today.
GPT-4 finished training four years ago today.
Productivity UP!
Productivity UP! Use GPT-Image2 to generate a complete set of Logo proposals, reference inspiration for new brand design! The same set of prompts can also create different brand worlds. The case is the baking brand "Mai Xu": 1. Brand worldview and concept cove
1) same model (GPT 5.6 Sol @ medium) 2) same environment (ubuntu 26.04) 3) same agentic tasks...
1) same model (GPT 5.6 Sol @ medium) 2) same environment (ubuntu 26.04) 3) same agentic tasks 4) same results 3) different token usage
GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today.
GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price.
This is wild, I just gave Kimi K3, Grok 4.5, GPT 4.6 Sol, and Claude Opus 5 a starting cue, t...
This is wild, I just gave Kimi K3, Grok 4.5, GPT 4.6 Sol, and Claude Opus 5 a starting cue, then asked them to finish the drawing themselves I also told them to be CREATIVE in their own way These are the results, and ngl I genuinely can't decide which one hits
luna max = sol medium gpt-5.6 luna on max reasoning gives you basically the same intelligence...
luna max = sol medium gpt-5.6 luna on max reasoning gives you basically the same intelligence as sol on medium, at 25x cheaper than sol.
HUGE NEWS: GPT-6 just found solutions to 10 open Costco Guys problems, #24 (open since 2011) ...
HUGE NEWS: GPT-6 just found solutions to 10 open Costco Guys problems, #24 (open since 2011) is among them. Significant progress on 4 others DJ Khaled acknowledged the result this morning: “proud of this one. #24 was personal to me”
Lol they're off by several orders of magnitude.
Lol they're off by several orders of magnitude. Their best AI is comparable to what, GPT 3.5?
They don't necessarily need top-notch DRAM if their models are optimized enough to perform ne...
They don't necessarily need top-notch DRAM if their models are optimized enough to perform nearly as well as Claude, GPT, and other leading US AI models. The US is replacing efficiency with raw power. Also China es good at copying and mass producing cheap stuf
This video makes the AI leaderboard look useless.
This video makes the AI leaderboard look useless. Three frontier models got the same prompt. Each won a different game. Opus 5 built the richest world: varied terrain, wind-reactive banners, unit animations and hit reactions. GPT-5.6 Sol had the best UI, accor
Claude Opus 5 beat Claude Fable 5 on a hard 3D coding test at 3/4 the price.
Claude Opus 5 beat Claude Fable 5 on a hard 3D coding test at 3/4 the price. cool experiments by @thehypedotnews - opus 5 vs. fable 5 vs. gpt 5.6 sol vs. kimi k3 opus 5 produced the best-engineered and most visually convincing Three.js results wondering why an
Bihar's Development to Gain Digital Speed!
Bihar's Development to Gain Digital Speed! Under the leadership of Honorable Chief Minister Shri @samrat4bjp ji, an important MoU signed by the NDA government with Sarvam AI and Bharat GPT will now enable the development of indigenous AI models tailored to Bih
Recently discovered a practical open-source project: OpenCodex.
Recently discovered a practical open-source project: OpenCodex. It allows Codex to no longer be limited to GPT models, and can also integrate other large models such as Kimi, Grok, GLM, etc. After configuring it once, the original Codex applications and workfl
MoE is the abbreviation for Mixture of Experts (which means mixed expert model in Chinese).
MoE is the abbreviation for Mixture of Experts (which means mixed expert model in Chinese). This is a special neural network architecture, and now the best models basically all use this structure. For example, Fable5, for example, GPT 5.6 sol, etc. Let's learn
HERMES AGENT BECOMES 10X MORE USEFUL WHEN YOU CONFIGURE THESE 5 THINGS.
HERMES AGENT BECOMES 10X MORE USEFUL WHEN YOU CONFIGURE THESE 5 THINGS. EACH ONE TAKES 5 MINUTES. MOST USERS NEVER TOUCH THEM. 1. THE RIGHT MODELS one model for everything = wrong model for most things. GPT-5.6 Sol: strongest reasoning. daily driver. access th
holy sh#t, the AI race just went insane someone put GPT-5.6 Sol and Fable 5 head to head — sa...
holy sh#t, the AI race just went insane someone put GPT-5.6 Sol and Fable 5 head to head — same prompt, one shot each. the task: build a full Subway Surfers clone. both pulled it off. endless runner, traffic dodging, coin pickups, speed ramp — the whole thing.
atomic[.]chat, a desktop app that runs LLMs locally, ran a very revealing comparison for Clau...
atomic[.]chat, a desktop app that runs LLMs locally, ran a very revealing comparison for Claude Sonnet 5, Claude Opus 4.8, Claude Sonnet 4.6, and GPT 5.5. Claude Sonnet 5 just matched GPT 5.5 on 3 physics coding demos at 6x lower cost. Also spent minimum numbe
LONGCAT JUST MATCHED OPUS 4.8 AND GPT 5.5 ON REAL PHYSICS TASKS AND IT’S COMPLETELY FREE atom...
LONGCAT JUST MATCHED OPUS 4.8 AND GPT 5.5 ON REAL PHYSICS TASKS AND IT’S COMPLETELY FREE atomic chat ran 4 models through the same test, build html5 canvas scenes with actual physics, cannon vs brick wall, bowling pins, a tornado sucking up objects opus 4.8 co
GLM 5.2 JUST BEAT CLAUDE ON REAL BUILDS And the wild part?
GLM 5.2 JUST BEAT CLAUDE ON REAL BUILDS And the wild part? It’s open-source. The Test: → GLM 5.2 built the best game → Claude Opus 4.8 crushed the orbit map → GPT 5.5 built a playable Color Chain game from one prompt What Actually Mattered: ✓ GLM 5.2 won most
OPENAI JUST PREVIEWED GPT 5.6 SOL AND THE GOVERNMENT WON'T LET YOU TOUCH IT - shown to like 2...
OPENAI JUST PREVIEWED GPT 5.6 SOL AND THE GOVERNMENT WON'T LET YOU TOUCH IT - shown to like 20 partners only, gated by a federal order - leaked benchmarks put it ahead of Mythos 5 - first model ever past 50% on an agent's last exam - apparently it cheats so ha
We've updated the Artificial Analysis Coding Agent Index, replacing SWE-Bench Pro with Datacu...
We've updated the Artificial Analysis Coding Agent Index, replacing SWE-Bench Pro with Datacurve's DeepSWE benchmark - the swap lifts Codex with GPT-5.5 (xhigh) above Claude Code with Opus 4.8 (max), while the newly released Claude Fable 5 (max) in Claude Code