Back to AI Pulse
TAG COLLECTION

\bGPT\b

All AI Pulse updates tagged "\bGPT\b".

34 signals
X.com08/22, 22:33Business

FOUR MODELS.

FOUR MODELS. THREE TASKS. ZERO FIXES. everything looked the same at the start. but after one prompt, the difference became obvious. GPT-5.6, Claude Opus 5, Qwen 3.8 Max and Grok 4.6 had to build a premium product page, a playable browser game, and an animated

X.com08/19, 00:03Research

I ran 10 prompts blind through GPT-4o, Claude 5, and Grok 4.6.

I ran 10 prompts blind through GPT-4o, Claude 5, and Grok 4.6. Same task. No labels. Just the output. Results: > 3 tasks: GPT-4o won > 4 tasks: Claude 5 won > 3 tasks: Grok 4.6 won 2-point spread on benchmarks. My test confirms it. The model stopped being the

X.com08/19, 00:03Infrastructure

Anthropic Engineer Andrej Karpathy dropped a full 6-hour guide on "Intro and Deep Dive to LLM...

Anthropic Engineer Andrej Karpathy dropped a full 6-hour guide on "Intro and Deep Dive to LLMs and Tokenizer: • 00:00 - Intro to LLMs • 00:59:28 - Deep dive into LLMs (Karpathy method) • 03:31:43 - Let s build the GPT Tokenizer from scratch This Karpathy cours

X.com08/17, 00:20Agent

GPT-5.6 Sol 1M in Codex.

GPT-5.6 Sol 1M in Codex. This used to only work for API keys, but we just flipped the switch and works for usage through ChatGPT accounts now too. The same warning applies, there is a reason the current context length is the default, we have tuned it to ~perfe

X.com08/17, 00:20Agent

Here is how to enable a 1M-token context window in Codex for GPT-5.6 Sol.

Here is how to enable a 1M-token context window in Codex for GPT-5.6 Sol. Even though we have tuned the context limit in Codex to be set optimally when it comes to performance and cost, this is a common ask, so here it is documented. A larger context window le

X.com08/12, 00:34Product Update

deepseek 马上发布code版本,可以直接调用工具跟执行 马上国内可以调用电脑的版本就出来了 按照deepseek的便宜价格,应该会抢很大一部分GPT 的市场 目前公众号已经注册了...

deepseek 马上发布code版本,可以直接调用工具跟执行 马上国内可以调用电脑的版本就出来了 按照deepseek的便宜价格,应该会抢很大一部分GPT 的市场 目前公众号已经注册了,马上开始招测试用户

X.com08/11, 19:10Agent

.

. @Toyota built their R&D GPT on Deep Agents to help employees search years of paint, seat, and manufacturing research with a single command. Here's their story.

X.com08/11, 00:32Safety

do you understand what OpenAI just did?

do you understand what OpenAI just did? built a hacking-grade AI, then locked it behind an approval process for “trusted defenders” only it’s called Daybreak, and the headline model is GPT-5.6-Cyber. the numbers tell the story. on advanced cybersecurity tasks,

X.com08/10, 21:47Business

Developers weigh in on OpenAI's GPT-5.6 Sol, from database audits to Erdős problem solving, a...

Developers weigh in on OpenAI's GPT-5.6 Sol, from database audits to Erdős problem solving, as it takes on Anthropic's Claude Opus 5.

X.com08/09, 19:27Infrastructure

Fable 5 is great at under-defined feature building GPT 5.6 Sol is great at most workhorse tas...

Fable 5 is great at under-defined feature building GPT 5.6 Sol is great at most workhorse tasks Grok 4.5 is great where token efficiency matters, like PR review Opus 5 GPT 5.6 Luna Max is 80% of Sol at way less cost

X.com08/09, 00:08Business

翻到一条六年前发的、体验 GPT-3 后的推文,没想到这一晃就六年了,这下真的开始慌了。

翻到一条六年前发的、体验 GPT-3 后的推文,没想到这一晃就六年了,这下真的开始慌了。

X.com08/08, 23:46Business

I just realized I will be in San Francisco again on Sep 30th and that this is the exact date ...

I just realized I will be in San Francisco again on Sep 30th and that this is the exact date I predicted I would turn more bullish on OpenAI than Anthropic maybe they delay GPT-6 until DevDay on Sep 29th, which could technically be Sep 30th in Germany (if that

X.com08/03, 00:47Business

Productivity UP!

Productivity UP! Use GPT-Image2 to generate a complete set of Logo proposals, reference inspiration for new brand design! The same set of prompts can also create different brand worlds. The case is the baking brand "Mai Xu": 1. Brand worldview and concept cove

X.com08/02, 23:46Agent

1) same model (GPT 5.6 Sol @ medium) 2) same environment (ubuntu 26.04) 3) same agentic tasks...

1) same model (GPT 5.6 Sol @ medium) 2) same environment (ubuntu 26.04) 3) same agentic tasks 4) same results 3) different token usage

X.com08/02, 23:31Infrastructure

GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today.

GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price.

X.com08/02, 23:02Business

This is wild, I just gave Kimi K3, Grok 4.5, GPT 4.6 Sol, and Claude Opus 5 a starting cue, t...

This is wild, I just gave Kimi K3, Grok 4.5, GPT 4.6 Sol, and Claude Opus 5 a starting cue, then asked them to finish the drawing themselves I also told them to be CREATIVE in their own way These are the results, and ngl I genuinely can't decide which one hits

X.com08/02, 22:46Business

luna max = sol medium gpt-5.6 luna on max reasoning gives you basically the same intelligence...

luna max = sol medium gpt-5.6 luna on max reasoning gives you basically the same intelligence as sol on medium, at 25x cheaper than sol.

X.com08/02, 22:17Business

HUGE NEWS: GPT-6 just found solutions to 10 open Costco Guys problems, #24 (open since 2011) ...

HUGE NEWS: GPT-6 just found solutions to 10 open Costco Guys problems, #24 (open since 2011) is among them. Significant progress on 4 others DJ Khaled acknowledged the result this morning: “proud of this one. #24 was personal to me”

X.com07/29, 20:46Business

They don't necessarily need top-notch DRAM if their models are optimized enough to perform ne...

They don't necessarily need top-notch DRAM if their models are optimized enough to perform nearly as well as Claude, GPT, and other leading US AI models. The US is replacing efficiency with raw power. Also China es good at copying and mass producing cheap stuf

X.com07/27, 00:03Business

This video makes the AI leaderboard look useless.

This video makes the AI leaderboard look useless. Three frontier models got the same prompt. Each won a different game. Opus 5 built the richest world: varied terrain, wind-reactive banners, unit animations and hit reactions. GPT-5.6 Sol had the best UI, accor

X.com07/25, 00:47Business

Claude Opus 5 beat Claude Fable 5 on a hard 3D coding test at 3/4 the price.

Claude Opus 5 beat Claude Fable 5 on a hard 3D coding test at 3/4 the price. cool experiments by @thehypedotnews - opus 5 vs. fable 5 vs. gpt 5.6 sol vs. kimi k3 opus 5 produced the best-engineered and most visually convincing Three.js results wondering why an

X.com07/22, 23:17Business

Bihar's Development to Gain Digital Speed!

Bihar's Development to Gain Digital Speed! Under the leadership of Honorable Chief Minister Shri @samrat4bjp ji, an important MoU signed by the NDA government with Sarvam AI and Bharat GPT will now enable the development of indigenous AI models tailored to Bih

X.com07/22, 22:47Agent

Recently discovered a practical open-source project: OpenCodex.

Recently discovered a practical open-source project: OpenCodex. It allows Codex to no longer be limited to GPT models, and can also integrate other large models such as Kimi, Grok, GLM, etc. After configuring it once, the original Codex applications and workfl

X.com07/22, 22:47Business

MoE is the abbreviation for Mixture of Experts (which means mixed expert model in Chinese).

MoE is the abbreviation for Mixture of Experts (which means mixed expert model in Chinese). This is a special neural network architecture, and now the best models basically all use this structure. For example, Fable5, for example, GPT 5.6 sol, etc. Let's learn

X.com07/22, 20:32Agent

HERMES AGENT BECOMES 10X MORE USEFUL WHEN YOU CONFIGURE THESE 5 THINGS.

HERMES AGENT BECOMES 10X MORE USEFUL WHEN YOU CONFIGURE THESE 5 THINGS. EACH ONE TAKES 5 MINUTES. MOST USERS NEVER TOUCH THEM. 1. THE RIGHT MODELS one model for everything = wrong model for most things. GPT-5.6 Sol: strongest reasoning. daily driver. access th

X.com07/19, 20:47Business

holy sh#t, the AI race just went insane someone put GPT-5.6 Sol and Fable 5 head to head — sa...

holy sh#t, the AI race just went insane someone put GPT-5.6 Sol and Fable 5 head to head — same prompt, one shot each. the task: build a full Subway Surfers clone. both pulled it off. endless runner, traffic dodging, coin pickups, speed ramp — the whole thing.

X.com07/01, 00:46Infrastructure

atomic[.]chat, a desktop app that runs LLMs locally, ran a very revealing comparison for Clau...

atomic[.]chat, a desktop app that runs LLMs locally, ran a very revealing comparison for Claude Sonnet 5, Claude Opus 4.8, Claude Sonnet 4.6, and GPT 5.5. Claude Sonnet 5 just matched GPT 5.5 on 3 physics coding demos at 6x lower cost. Also spent minimum numbe

X.com06/30, 23:16Business

LONGCAT JUST MATCHED OPUS 4.8 AND GPT 5.5 ON REAL PHYSICS TASKS AND IT’S COMPLETELY FREE atom...

LONGCAT JUST MATCHED OPUS 4.8 AND GPT 5.5 ON REAL PHYSICS TASKS AND IT’S COMPLETELY FREE atomic chat ran 4 models through the same test, build html5 canvas scenes with actual physics, cannon vs brick wall, bowling pins, a tornado sucking up objects opus 4.8 co

X.com06/29, 21:31Open Source

GLM 5.2 JUST BEAT CLAUDE ON REAL BUILDS And the wild part?

GLM 5.2 JUST BEAT CLAUDE ON REAL BUILDS And the wild part? It’s open-source. The Test: → GLM 5.2 built the best game → Claude Opus 4.8 crushed the orbit map → GPT 5.5 built a playable Color Chain game from one prompt What Actually Mattered: ✓ GLM 5.2 won most

X.com06/29, 19:15Agent

OPENAI JUST PREVIEWED GPT 5.6 SOL AND THE GOVERNMENT WON'T LET YOU TOUCH IT - shown to like 2...

OPENAI JUST PREVIEWED GPT 5.6 SOL AND THE GOVERNMENT WON'T LET YOU TOUCH IT - shown to like 20 partners only, gated by a federal order - leaked benchmarks put it ahead of Mythos 5 - first model ever past 50% on an agent's last exam - apparently it cheats so ha

X.com06/14, 20:15Agent

We've updated the Artificial Analysis Coding Agent Index, replacing SWE-Bench Pro with Datacu...

We've updated the Artificial Analysis Coding Agent Index, replacing SWE-Bench Pro with Datacurve's DeepSWE benchmark - the swap lifts Codex with GPT-5.5 (xhigh) above Claude Code with Opus 4.8 (max), while the newly released Claude Fable 5 (max) in Claude Code