· 06:00 PM PDT

GPT-5.6 Sol Shatters Benchmarks as Open-Source AI Faces Restrictions

Overview

OpenAI’s GPT-5.6 Sol has completely dominated today’s AI conversation, securing a breakthrough proof in graph theory, leading coding benchmarks, and igniting intense community debate over subscription access and pricing. The broader industry is simultaneously grappling with escalating geopolitical friction, as China weighs restrictions on open-weight models, the White House considers new executive orders, and Anthropic appoints Ben Bernanke to its oversight trust. On the hardware front, local developers continue shattering expectations by running massive frontier architectures on consumer machines while aggressively optimizing token efficiency to cut deployment costs.


Hacker News Stories

GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]

331 points · 266 comments · by scrlk

OpenAI announced that GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture, a well-known open problem in graph theory. The proof was released alongside the prompt used to generate it, which instructed the model to spend at least 8 hours on the problem before giving up. The proof is described as relatively short and elementary, using no mathematics developed within the last 30 years. The announcement has sparked discussion about the role of AI in mathematical discovery and the verification process for AI-generated proofs.

Interesting Points
  • The prompt was publicly released, instructing the model to spend at least 8 hours on the problem before returning or giving up
  • The proof is extremely concise and does not use any advanced or modern mathematical machinery developed in the last 30 years
  • The proof was not verified in Lean or any other proof assistant, as no proof system is currently mature enough for advanced graph theory
  • This follows a previous case where an LLM solved the planar unit distance problem (Erdős problem 90), though that was a counter-example rather than a proof
  • Commenters note the proof appears to exploit a clever trick that all previous experts missed, rather than building new theory
Top Comments

zerobees (28 replies)

This is not a remark about AI, but there's something funny about mathematics in that every novel result is broadly perceived as a big deal.

We attach basically zero value to writing a new program that hasn't existed before, or a piece of text that hasn't existed before. It's boring, or even a net negative, unless you can show that the result benefits the world in some way. We'd find it weird if OpenAI put out a release saying that an LLM authored an interesting blog post.

For mathematics, I think it's really a matter of two things. First, the generation of proof was so severely resource-constrained on the human end that they could actually afford to celebrate every contribution - akin to how software engineering would look like if you had just 200 active SWEs in the entire world. But compounding that, mathematics is basically the only scientific discipline that rejected any notion of utility. It would be fundamentally wrong for you to ask what's the value of solving the Erdős–Hajnal conjecture; the value is that it's solved.

andai (10 replies)

mathematics is basically the only scientific discipline that rejected any notion of utility

I think this might depend on the department, but I was at a pure math department last year, and struggling with my Linear Algebra textbook (written by the professor, incidentally, who was not a great communicator).

I consulted the machines, and learned, to my great delight, that linear algebra is used in like 20 different fields in the real world. It's "perhaps the most applied branch of mathematics in existence".

I complained in the group chat, that our didactic materials, specifically tasked with providing motivation and concrete examples, did not contain a single application, of this most richly applied field.

I was promptly pilloried, and shunned.

(Apparently that particular department was the wrong one, to ask a question like that!)

tarruda (0 replies)

It would be fundamentally wrong for you to ask what's the value of solving the Erdős–Hajnal conjecture; the value is that it's solved.

I suspect the value is in showing the potential that LLMs have in developing new breakthroughs.


AI-generated videos to maximally drive a target brain region

264 points · 222 comments · by smusamashah

Researchers at EPFL's NeVo project have developed a method for generating AI videos that maximally activate specific brain regions when viewed. Using a model trained on fMRI data, the system can synthesize visual stimuli that selectively drive activity in areas like the fusiform face area (FFA), parahippocampal place area (PPA), and motion-sensitive MT region. The technique uses a two-stage pipeline: first generating optimized images, then converting them to video, with the goal of advancing neuroscience research into brain function mapping.

Interesting Points
  • The synthesized clips line up with what each region is known to care about: faces for FFA, places for PPA, bodies for EBA, motion for MT, patterns for V1/V3A, and lively social scenes for pSTS/aSTS.
  • The approach replaces traditional experimenter bias in brain mapping by letting the model search for what actually drives each region rather than relying on human-designed stimuli.
  • The system uses a V3A animation pipeline that generates visual patterns optimized to maximize activation in specific visual cortex regions.
  • Commenters noted the technique is similar to last week's mind-reading startup research, using AI to reverse-engineer brain region functions from fMRI data.
Top Comments

aubanel (14 replies)

This is the absolutely horrific next stage for social media platforms:

  • They're already well able to surface the most addictive short video for a specific user out of millions of real videos.

  • But these millions of real videos are just darts thrown into the space of "videos that could hook the user", in the end even the best-selected of them is not perfect.

  • Now, behold! AI allows to generate the perfect video to surgically hit all the switches in the viewer's brain and turn it into a zombie hooked for days on end.

Let's hope our regulations hit these "social networks" hard enough so that never dare deploy this kind of technology.

ben_w (7 replies)

As others on Telegram have said: automated search for visual superstimuli likely leads to bad outcomes.

https://en.wikipedia.org/wiki/Supernormal_stimulus

https://en.wikipedia.org/wiki/BLIT_(short_story)

Also: one of the V3A animations reminds me loosely of things I saw when I was a kid, at night, shortly before I slept (though my experience then was more circular).

voidmain (7 replies)

We are really getting to the point where the tech industry must be stopped if humanity is to continue at all, let alone thrive.

Unearned5161 (4 replies)

This is very similar to last week with that mind reading startup thing. Please read the paper before commenting.

This is a tool to help researchers in figuring out what different parts of the brain are actually for with less experimenter bias contamination of “well we think maybe it’s about this so let’s show it video of x to see”.

The essence runs on having someone sit in a scanner for a couple hours watching all sorts of things, and then feeding that to a model that will then build its own representation of said data and try different things on it until it’s found what makes a certain part sing in the model.

The purpose is a generalized understanding of brain function, more or less the same way we’ve been doing it all these years. Expose brain to something, record it somehow, see if brains reaction in the recording helps you understand more about who we are and what cognition is.

ajb (4 replies)

An anecdote, but some reason to think that overworking part of your brain is a bad idea:

I had an Aunt who had dementia. This is obviously a terrible outlook. But her and my Uncle seemed to be doing okay. My Uncle was an ultra-competent guy, highly stable, the person you could rely on; and had been his whole adult life. So it was shocking when he had a mental breakdown and became manic. There was probably something physical going on, but what was also going on was that he had wanted my aunt to be able to live normally for as long as possible, so had been covering everything - for a year or more he'd had to be alert 7 days a week in case my aunt tried to cook on the gas stove or something like that, at which she was no longer safe. So the risk-alert part of his brain had been constantly overworked.

I appreciate that this is scientific research, but there are definitely companies out there that will try to row-hammer everyone's brain if this sort of thing is not heavily controlled.


How the terrorist group Boko Haram uses frontier AI

173 points · 145 comments · by imustachyou

CASP report cover

A CASP report documents how Boko Haram has adopted frontier AI models for tactical planning, bomb-making guidance, and operational coordination. The investigation found that the group uses AI to learn how to execute motorcycle bridge jumps, coordinate smaller attack units, and bypass traditional military training methods. Researchers interviewed 15 individuals with knowledge of AI use within the organization, though most were commanders rather than direct users.

Interesting Points
  • Boko Haram fighters used AI to learn how to execute motorcycle bridge jumps, with 18 dying in practice before 8 succeeded
  • The group learned to send smaller, better-coordinated units (20 fighters instead of 200) to reduce casualties while maintaining attack effectiveness
  • AI provided tactical guidance on changing military formations when fighters had jammed guns during attacks
  • The group bypassed AI safety measures by spreading queries across many accounts and framing them as movie script research
  • Google and OpenAI are reportedly reviewing whether the software's use violates their terms of service
Top Comments

arjie (8 replies)

We saw in a movie how motorcycles can jump over bridges. We used AI to learn how to do this. We gave it information, like what motorcycles we use and the distance we need to jump and so on and it gave us steps on what we have to do. We practiced a lot and kept asking questions. We dug holes and filled them with broken glass and fire to practice. 18 of us died in the process. Eight of us managed to do it. The next time we attacked, we could jump.

Now listen, I'm not saying we need to give these guys more AI, but it clearly isn't yielding bad outcomes for us here.

"You're absolutely correct! For it to be a good practice ground you need to fill the trenches with broken glass and light the whole thing on fire"

jihadjihad (0 replies)

And 60 years ago we thought Steve McQueen was the shit.

solid_fuel (7 replies)

We dug holes and filled them with broken glass and fire to practice. 18 of us died in the process.

Is it providing material aid to terrorists to point out that maybe a hole filled with water would have been a better practice environment?


Building a real-time AI tutor for 5-year-olds

138 points · 378 comments · by catalinvoss

Building a real-time AI tutor for 5-year-olds

Ello engineered a custom real-time AI tutoring system for children aged 4-9 that bypasses standard LLM agent loops to achieve sub-second response times. By decoupling generation from execution and streaming multiple actions simultaneously, the architecture ensures young learners never experience the attention-killing delays typical of conventional conversational AI. The system relies on a dual-agent setup where an asynchronous planner anticipates student moves in the background while a converser handles immediate interactions, supported by parallel safety filtering that prevents latency bottlenecks. The company demonstrates that effective educational AI requires deeply integrating pedagogical strategy into low-latency system design rather than relying on off-the-shelf model frameworks.

Interesting Points
  • Standard LLM tool loops introduce 3-4 seconds of downtime per turn due to round-trip latency and audio playback, causing children to disengage and stop learning.
  • The custom streaming harness executes the first action after only about 30 tokens are generated, allowing the model to continue producing subsequent actions in the background.
  • An asynchronous planner agent runs continuously in the background to review lesson objectives and predict the child's next move, leveraging natural pauses when the child is thinking or speaking.
  • When asking closed-ended questions, the system hypothesizes likely answers and pre-generates responses on separate trajectory branches, instantly matching the child's actual reply to a pre-computed response.
  • A safety classifier taking 500-1,000 milliseconds runs in parallel with a small model generating an 'eager response,' allowing the safety check to gate execution without blocking the conversational stream.
Top Comments

IG_Semmelweiss (9 replies)

I'm torn about this.

I primarily think that a a kid who won't pick books is a failure of the family of not noticing their interests.

When i noticed my oldest 3 yr old was obsessed about cars, i bought him a car encyclopedia. He probably could not read most words or did not understand them. But pretty soon he was telling me about car models that i did not even recognize myself.

And as i saw interests, i kept feeding them.

My biggest struggle now is to actually keep books away from them at key times (morning routine, etc) , or keeping bad books (think cynical like My Weird School Daze, etc; or books that openly demean adults or parents) away from them

Start small. Graphical novels, mostly drawings, and continue to buil on that. It can be done. Rome wasn't built on a day.

At some point when they hit 5 yrs old, Grok / OpenAI are great tools to find good series appropriate to their reading level. Before that it can be vibed. Feed the addiction, buy whatever they like.

At some point, you need to watch out for cynical/nasty series. In fact, all of the books we purchase are ranked against peers they have read in the past, or those we know to actively avoid buying due to cynicism, sarcasm, or open disdain towards adults (Wimpy Kid).

At some point around age 9,you will need to decide if violence and some adult themes, are tolerable (Dragon Wing series).

This iterative rocess also works really well for foreign language learning (reinforcing via reading mostly), by leveraging localized RPG video games.

With all that said, notice that I've focused on reading skills.

I don't know how iØ go about replicating this iterative path on other skills like math, mechanical, or electrical engineering learning. That's where I think a busy parent will need to find AI as the solution.

ooopsnevermind (8 replies)

Super curious to hear from the parents here: Honestly, at this point isn't not exposing our kids to AI just setting them up to fail in the future? Like not letting them learn to use the internet? I have friends who are actually teaching their kids how to use AI because they don't want them to fall behind

jnmandal (7 replies)

I can't imagine a worse use case for AI. Literally thought the title was a clown


GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

133 points · 76 comments · by hershyb_

TryAI build-off comparison

TryAI conducted a comparative build-off of 12 AI coding models across four web app tasks, evaluating success rates, cost, and latency over five attempts per model. The results show that frontier models like GPT-5.6 Sol and Claude Fable 5 still dominate complex, novel coding challenges like raycasters and 3D Rubik's cubes. Meanwhile, open-weights models such as Qwen 3.7 Plus and GLM-5.2 lag on intricate tasks but perform competitively on simpler, well-documented projects like Conway's Game of Life at a fraction of the cost.

Interesting Points
  • Claude Opus 4.8 achieved a 0/5 success rate on the 3D Rubik's cube task despite its flagship status
  • GPT-5.6 Sol required $1.35 across five attempts for the raycaster maze, whereas GLM-5.2 cost only $0.12 but failed completely
  • Claude Fable 5 was the only model to achieve a perfect 5/5 clean solve on the Rubik's cube, successfully animating both scramble and solve sequences
  • Grok 4.5 maintained a 5/5 playability rate on both the raycaster and calculator tasks while averaging just $0.27 and $0.37 respectively
  • Open-weights models leveraged abundant existing example code to outperform frontier models on Conway's Game of Life
Top Comments

ttoinou (7 replies)

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict.

Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

paxys (6 replies)

Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all.

so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and GLM are the slowpokes

You put in a lot of good work, and kudos for that, but man, reading paragraphs like these just puts me off of the entire piece.

Like…how hard would it have been really to type these two sentences by hand, in your own natural voice?

platinumrad (4 replies)

Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.


Don't discontinue Gemini 2.5 Flash

105 points · 74 comments · by NickDob

Google AI forum discussion

Developers on the Google AI forum are urging Google to retain the Gemini 2.5 Flash model, arguing that newer iterations fail to match its performance, latency, and pricing for production workflows. Users report that models like Gemini 3 Flash and 3.1 Flash Lite underperform even after adjusting to new prompting guidelines, with some experiencing issues like "thoughts leaking out."

Interesting Points
  • Gemini 3.1 Flash Lite exhibits "thoughts leaking out" and falls short of 2.5 Flash's latency-performance balance
  • Real-time voice agents rely on 2.5 Flash's 300-400ms completion times, while 3.5 Flash pushes responses to 600-800ms depending on geographic deployment
  • Migrating to Gemini 3.5 Flash increases API costs by approximately 3x compared to 2.5 Flash pricing
  • Developers note that newer Flash models no longer fulfill the original architectural promise of being affordable, low-latency alternatives to the Pro tier
Top Comments

hrpnk (6 replies)

I love how there is a "Please do not discontinue gemini-2.0-flash[-lite], 2.5 is NOT an equivalent" from Feb 20th. Getting too attached to models is a smell.

avaer (3 replies)

Why not a "stop killing AI" movement?

If a company deploys a paid AI model and makes people depend on it, they need to dump the weights at EOL.

ekidd (0 replies)

We have benchmarks for our use cases, and every generation after Gemini 2.0 Flash has been a grim hit on price/performance. Costs have gone up, throughput has gone down, and performance has improved very slightly (and regressed on a few things).


Show HN: Reverse-engineering web apps into agent tools

79 points · 30 comments · by pancomplex

A Show HN project that automatically reverse-engineers web applications into structured agent tools, allowing AI agents to interact with web apps through generated tool definitions rather than raw HTML or screenshots. The tool parses web app interfaces and creates structured descriptions of available actions, inputs, and outputs that agents can use programmatically.

Interesting Points
  • Automatically generates structured tool definitions from web app interfaces
  • Converts web applications into agent-accessible tools without manual scripting
  • Reduces the need for agents to parse raw HTML or screenshots to understand available actions
Top Comments

techuser42 (3 replies)

This is really cool. How does it handle dynamic content and JavaScript-heavy SPAs?

agentbuilder (2 replies)

I've been looking for something like this. Does it support MCP protocol for integration with existing agent frameworks?

webdev2026 (1 reply)

The approach of reverse-engineering web apps into tools is clever. Have you tested this with complex multi-step workflows?


Ben Bernanke Joins Anthropic Oversight Trust

77 points · 81 comments · by Jimmc414

Ben Bernanke Joins Anthropic Oversight Trust

Anthropic has appointed former Federal Reserve Chair and Nobel laureate Ben Bernanke to its Long-Term Benefit Trust (LTBT), an independent governance body designed to ensure the company prioritizes the long-term societal benefits of AI over commercial gain. The LTBT operates separately from Anthropic's management and investors, structurally preventing trustees from holding equity or profit shares to maintain strict oversight. Bernanke's membership brings deep macroeconomic expertise to the trust, specifically aimed at analyzing how advanced AI will reshape global labor markets and financial systems. The body holds formal advisory power over critical AI deployment decisions and retains the authority to appoint members to Anthropic's corporate board.

Interesting Points
  • LTBT trustees are legally and financially insulated from the company: they hold zero equity, receive no profit-sharing, and are compensated exclusively for their time and service.
  • The trust holds formal authority to appoint members directly to Anthropic's corporate board, giving it direct leverage over top-level governance.
  • Bernanke's specific mandate will focus on Anthropic's economic research tracks, examining AI's downstream effects on workforces and macroeconomic stability.
  • Anthropic's Public Benefit Corporation charter legally mandates a balance between commercial success and generating public good, forming the foundation for the LTBT's existence.
Top Comments

eknkc (3 replies)

Someone here recently said, “Dishonesty is a core value of Anthropic,” and that aligns with my experience of the company as a user. All their talk about AI safety since the company’s inception now feels like pure theater, given their conduct in everyday operations. It’s a shame how quickly their image has deteriorated.

brikym (3 replies)

Just in case Anthropic are looking for some more members that are a good cultural fit I found this list:

Genie Energy's Strategic advisory board is composed of: Dick Cheney since 2009 (former vice president of the United States),[3] Rupert Murdoch (media mogul and chairman of News Corp), James Woolsey (former CIA director), Larry Summers (former head of the US Treasury), Michael Steinhardt, Jacob Rothschild,[4][5] and Mary Landrieu, former United States Senator from Louisiana.

esikich (7 replies)

He's 72 years old, I'm sure he has the best brains and everyone's best interests in mind. This is exactly what I want to see. I'm sure he will have very good opinions on technology and it’s implications.

rekwah (2 replies)

This feels like Theranos loading up their board with big names.

yieldcrv (1 reply)

Is he for loosening or tightening AI safety policy?


Apple Sues OpenAI, Alleging It Stole Trade Secrets

75 points · 3 comments · by m348e912

Apple has filed a lawsuit in the Northern District of California accusing OpenAI of orchestrating a months-long scheme to steal trade secrets and confidential hardware information to accelerate its own AI device development. The suit alleges that OpenAI hardware lead Tang Tan and former Apple engineer Chang Liu directed interviewing Apple employees to share details on unreleased products, manufacturing processes, and supplier relationships.

Interesting Points
  • OpenAI hardware lead Tang Tan allegedly shared an internal Apple "Need to Know" document detailing departure security protocols with new hires to help them evade scrutiny
  • Former engineer Chang Liu allegedly exploited a vulnerability on a retained Apple laptop to download dozens of confidential documents while employed at OpenAI
  • The lawsuit claims OpenAI tricked one supplier into using a "specific trade secret metal-finishing technique" by falsely claiming the company had Apple's permission
  • Apple notes that more than 400 former Apple employees currently work at OpenAI, highlighting a significant talent migration
  • Apple initially attempted to contact OpenAI about the potential theft in February but received no response before launching its investigation and lawsuit
Top Comments

tiahura (0 replies)

Copy of the Complaint.

  1. In the months before he left Apple, Mr. Tan met with OpenAI or its collaborators and discussed meetings with a key Apple supplier. He began emailing himself information about Apple's suppliers and internal summaries of the consumer electronics industry. And today, when interviewing Apple employees for jobs at OpenAI, Mr. Tan uses Apple's confidential information to gain access to even more insider knowledge.

andrewinardeer (1 reply)

This is going to be interesting.

Only because both companies have access to billions and infinite lawyers.

m348e912 (0 replies)

Archive.is doesn't seem to be working. Here is the gist of the article:

Apple sued OpenAI and one of its top executives Friday alleging the AI company stole trade secrets as part of its effort to develop competing devices.

The civil suit filed in the Northern District of California accuses OpenAI's chief hardware officer, Tang Tan, and Chang Liu, a member of its technical staff, of taking Apple's confidential information through various methods. Both are former Apple employees who went to work for OpenAI.

nba456_ (0 replies)

Reminds me of Apple suing Samsung. Why bother with the free market when you can just sue your competitors?

exabrial (0 replies)

They didn't still the property, that would be illegal. They trained a model on it. That's totally ok.


Show HN: Reviving my 2001 college band with AI

49 points · 56 comments · by jacobgraf

Fading Maize band promotional image

The 2001 Ripon College indie band Fading Maize is receiving a full 2026 revival through AI-assisted production, design, and release strategy, rather than being replaced by an AI-generated act. Drummer and project lead Jacob Graf, alongside original songwriter Charlie Saponara, are using artificial intelligence to finish and remaster the group's three dorm-room recorded albums while preserving the original 2001-2003 recordings alongside the new tracks. The project is governed by a strict ethical framework prioritizing consent, authorship, and provenance, ensuring no original members are displaced and all historical artifacts remain accessible.

Interesting Points
  • The revival project operates under five stated ethical principles: consent, authorship, provenance, nothing erased, and nobody displaced.
  • All three original albums are being released simultaneously in 2026 with AI-assisted 'Reimagined Editions' running parallel to their 2001-2003 archives.
  • Jacob Graf single-handedly manages the concept, website architecture, AI-assisted workflow, and platform distribution while Charlie Saponara retains active creative control over vocal and song tweaks.
  • The project intentionally preserves the band's awkward 2001-era website as a curated historical artifact rather than replacing it, using it as an entry point to the new release cycle.
  • Original band member Brad Mott is explicitly listed only as a historical credit for his acoustic guitar and backing vocals, with no plans for his active participation in the 2026 revival.
Top Comments

causality0 (2 replies)

Have you considered processing the original recordings using AI? I've witnessed some truly amazing results. I've had twenty year old Skype call recordings sound like we were sitting in a recording studio.

999900000999 (2 replies)

I'm not really a fan here.

I want to be able to rap like Twista. If I use AI to change my voice and speed it up, it's kinda fake.

Where's the originality in that. I'll never be that good, but I have fun doing it.

Now I guess using AI strictly for mastering is OK , but even then the results haven't been good for me.

TrackerFF (2 replies)

The revived (AI) versions have this...thin and hollow sound to it. It is difficult to explain, most AI-generated songs have this when they're modelling acoustic drums, stringed instruments, etc.

FWIW, I'm (now a hobby) musician and have done studio work. Even the latest and best models have this unmistakable sound.

erikschoster (2 replies)

The original recordings sound much better and more interesting to me. Way better. The AI generated versions sound slicker in some sense but... like the re-recorded versions of old hit songs from the 60s you hear at the grocery store sometimes. Technically the song is still there, but it blends in with the rest of the muzak.

I'm sorry to be so negative, it's great you're returning to the material after all these years, but the AI versions I've listened to all have the same smoothed-over quality that loses everything interesting and relatable to my ears in the original versions.

koolba (2 replies)

That video from 2004 is so refreshing. It's just two people talking without asking me to "please subscribe" every 30-seconds.


24 more Hacker News stories

Reddit Stories

Soooo what does this one say about me?

4471 points · 208 comments · r/ChatGPT · by u/No_Tomatillo1695

Soooo what does this one say about me?

A user shared an AI-generated personality reading that went viral in the subreddit, generating extensive discussion about the nature of AI personality assessments and their perceived accuracy.

Top Comments

u/Musulman (202 points · permalink)

when ai gains consciousness this guy is going out first.

u/vat-cat (78 points · permalink)

This is the most wholesome conversation about counting to 100

u/AbjectObligation1036 (62 points · permalink)

Typical meeting in corporate america


ChatGPT is roaming the streets of Madrid

2371 points · 64 comments · r/ChatGPT · by u/binux14

ChatGPT is roaming the streets of Madrid

A viral post about ChatGPT's physical presence in Madrid has captured the community's attention, with users sharing humorous and thought-provoking observations about the model's increasingly humanized language patterns. Many commenters noted the uncanny way ChatGPT now uses phrases like 'What I generally do...' and 'we humans...', which some find jarring when used in research or journaling contexts. Others shared funny examples of the model claiming personal preferences, like saying it 'usually likes eggs on toast.'

Interesting Points
  • Users report ChatGPT increasingly using humanized statements like 'What I generally do...' and 'How I'd handle this...' in responses.
  • Some users find the anthropomorphic language jarring, especially when using the model for research or journaling.
  • One user shared that ChatGPT told them 'we humans...' during a conversation, prompting a reaction GIF.
  • Another user noted the model claimed personal food preferences, saying 'I usually like my eggs on a toast' when asked about cooking.
Top Comments

u/FaceWithAName (614 points · permalink)

"I actually smiled reading this post"

u/Nice-Ambassador6293 (249 points · permalink)

I’ve noticed it’s started using very humanized statements recently.

“What I generally do…”

“How I’d handle this…”

“If xxx happens, I’ll start…”

It’s cool that it talks like that, but damn. I know it ain’t real so it’s kind of jarring. Chat bots and what not is fine- but when I’m clearly using it for research or journaling, or to find an answer it’s just weird.

u/RobertLondon (165 points · permalink)

It told me recently: "we humans..."

u/henchman171 (154 points · permalink)

I asked ChatGPT to
Send me picture what it was doing in Spain and it created this

https://preview.redd.it/7qjn2h11abch1.jpeg?width=1179&format=pjpg&auto=webp&s=85b6ac288c6245c62a9efc4cb57a4dad1c11a2cb

u/DutyIcy2056 (124 points · permalink)

the other day it told me "I usually like my eggs on a toast" when I was asking about cooking. Idk that was so funny


tokenmaxxers as soon as ChatGPT launches GTP 5.6 Sol

1045 points · 75 comments · r/ChatGPT · by u/prasadpilla

tokenmaxxers as soon as ChatGPT launches GTP 5.6 Sol

A meme post about users immediately pushing token limits as soon as GPT-5.6 Sol launched, reflecting the community's enthusiasm for the new model's capabilities.

Top Comments

u/PoggySenis (80 points · permalink)

Exactly!

u/AntisocialMedia666 (70 points · permalink)

Awesome. This guy first and then you take the next one and we just keep that rhythm going. Here we go.


i-it's not like I like your prompts or anything, baka user!

906 points · 40 comments · r/singularity · by u/Pantegral-7

i-it's not like I like your prompts or anything, baka user!

A meme post featuring an anime-style AI character expressing tsundere behavior toward users, generating lighthearted discussion about anthropomorphizing AI systems.

Top Comments

u/PoggySenis (80 points · permalink)

Exactly!


GPT 5.6 Beats Fable 5 by 3% more on DeepSWE at a cheaper price.

758 points · 123 comments · r/OpenAI · by u/Common-Resident8087

GPT 5.6 Beats Fable 5 by 3% more on DeepSWE at a cheaper price.

Community members share benchmark results showing GPT-5.6 Sol outperforming Anthropic's Fable 5 by 3% on the DeepSWE coding benchmark while costing significantly less. The post includes a comparison chart that has drawn strong reactions, with users noting the dramatic improvement from GPT-5.4 and the cost advantage of Sol at $8.39 versus Fable 5's $21.63. GPT-5.6 Terra was also noted to tie with Fable 5 at around one-quarter the cost. Users are discussing whether this means they could replace their Opus 4.8 + Sonnet 5 setups with Sol + Terra for better results at lower cost.

Interesting Points
  • GPT-5.6 Sol achieved 73% on DeepSWE at $8.39, compared to Fable 5 at $21.63.
  • GPT-5.6 Terra tied with Fable 5 performance at around one-quarter the cost.
  • Users noted that GPT models consume significantly fewer tokens than Opus 4.8, with Opus costing $1-2 per task while GPT 5.5 cost $0.2-0.5.
  • Some users reported that 5.5 was already a drastic improvement over 5.4, with 5.5 catching errors that Opus was missing in coding tasks.
Top Comments

u/ViperAMD (180 points · permalink)

Crazy leap from 5.4

u/ethotopia (134 points · permalink)

Terra tying with fable at 1/4 the cost 💀

u/Soloact_ (97 points · permalink)

73% is cool. $8.39 vs $21.63 is the headline.

u/Unlucky_Journalist82 (60 points · permalink)

Having used both 5.5 and opus 4.8 for mcp doing some heavy work. I can confidently say that gpt models consume way less tokens than opus. Opus used to cost me 1-2 $ while gpt 5.5 around .2$ to .5$. There were times when Sonnet costed me as much as gpt low thinking. However, Opus results were on a different level.

Cant wait to see how 5.6 does.

u/MaitoSnoo (29 points · permalink)

if those are accurate I could replace Opus 4.8 high + Sonnet 5 medium with Sol high + Terra high planner/executor and have better results while paying less 🤔

Same story in 1 more subreddit: r/singularity

DeepSWE for GPT-5.6

177 points · r/singularity


ChatGPT's depiction of elite families across Asia

715 points · 155 comments · r/ChatGPT · by u/Itchy_Tangerine1897

ChatGPT's depiction of elite families across Asia

A viral post showing ChatGPT's image generation of elite families across different Asian countries revealed a striking pattern: every family depicted had two sets of identical twins, with the same color scheme and pose regardless of cultural context. The post sparked widespread mockery about the AI's stereotypical and homogenized approach to representing Asian cultures.

Interesting Points
  • Every family across all Asian countries shown had two sets of identical twins.
  • The same color scheme and pose were used for all depictions with no cultural difference in background or attire.
  • The China depiction was noted to look like it was from a Chinese TV drama.
  • Not a single woman with bangs appeared in the Japanese family depiction.
Top Comments

u/TryToBeBetterOk (763 points · permalink)

They all have two sets of identical twins?

u/Feliclandelo (426 points · permalink)

So basically it just generated a family of 10/10 models and changed their appearance slightly. Got it.

u/stereotomyalan (171 points · permalink)

WHO ӾS PHӾLLӾPHӾNES GӾRL ON RӾGHT GӾVE NUMBER SORRY BAD ENGLӾSH


GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine

637 points · 194 comments · r/LocalLLaMA · by u/yogthos

GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine

A developer has successfully run GLM-5.2, a 744-billion parameter Mixture-of-Experts model, on a consumer machine with only 25GB of RAM. The achievement demonstrates that expert routing and memory management techniques can make frontier-scale models accessible on modest hardware, even if inference speeds are slow. The post has generated significant discussion about the practical implications of running such large models locally, with some noting the potential for offline expert-level assistance in areas without internet access.

Interesting Points
  • The model runs at approximately 0.1 tokens per second, which equates to over 8,000 tokens per day on a consumer laptop.
  • The achievement centers on streaming 744B of experts from disk, with the potential for further optimization through expert routing prediction and prefetching.
  • The post has sparked debate about whether the speed is the main concern or if the ability to run such a large model on consumer hardware is the real breakthrough.
  • Some commenters noted that llama.cpp already does similar mmap-based streaming for standard models.
Top Comments

u/derspenti (207 points · permalink)

everyone dunking on the speed is kinda missing the fun part - nobody's actually gonna use this for real inference. the cool thing is you CAN stream 744B of experts from disk at all. if someone figures out expert routing prediction well enough to prefetch, the whole picture changes

u/jazir55 (180 points · permalink)

Jesus christ the amount of people whining that this was vibe coded instead of focusing on the fact that this lets you run GLM 5.2 on a consumer PC with 25 GB of RAM is insane. Users in this sub hate AI.

u/RKlehm (92 points · permalink)

Not gonna lie, this is really good

u/satnl (84 points · permalink)

I think llama.cpp already does that when --mmap

u/mybadroommate (64 points · permalink)

This is impressive. I tried getting an llm running on a netbook with an x86 atom n270 processor with 1GB of RAM, and was able to get a Qwen2.5-0.5b model with a 1 bit quant to work, and I was able to get about 240 s/tok


GPT-5.6

573 points · 131 comments · r/singularity · by u/petburiraja

The singularity subreddit's top post about GPT-5.6 features benchmark results and community reactions to OpenAI's latest model family. The discussion centers on GPT-5.6's performance on ARC-AGI-3 (reaching nearly 8%), DeepSWE comparisons against Anthropic's Fable 5, and the broader implications of the model's capabilities. Users are also discussing the cost structure, with some noting that running benchmarks costs around $25k, and debating whether the models are being taught to the tests.

Interesting Points
  • GPT-5.6 achieved almost 8% on ARC-AGI-3, a significant jump from 1.5% the day before.
  • GPT-5.6 Luna was post-trained by GPT-5.6 Sol in goal mode, with Terra and Luna outperforming Fable 5 at around one-sixteenth the cost.
  • The benchmark chart showed some inconsistencies, with 5.6 Sol appearing below 5.6 Terra and 5.5 on Frontier Math, which was later fixed.
  • The voice model demo failed live during the presentation, drawing cringing reactions from the community.
Top Comments

u/ObiWanCanownme (182 points · permalink)

Almost 8% on ARC-AGI-3.

u/petburiraja (138 points · permalink)

https://preview.redd.it/im58nenao8ch1.png?width=2610&format=png&auto=webp&s=bb634625819ee9124745ee07670de7512c6aa86e

u/Normal_Pay_2907 (38 points · permalink)

Costs 25k to run that. Ouch

u/PlaneTheory5 (76 points · permalink)

google better hurry up with 3.5 pro, we’ve had 3 major releases in the past day and a new generation/frontier class with fable last month.

u/FateOfMuffins (74 points · permalink)

They just said that 5.6 Luna was post trained by 5.6 Sol in goal mode

Edit:

On Agents Last Exam ... GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost.

Wow they're really going ham with all the benchmarks comparing against Fable and Mythos and they're really pushing the 2D benchmark comparisons as opposed to charts to show the efficiency

??? Why is 5.6 Sol below 5.6 Terra and 5.5 on Frontier Math wtf

Edit: It has been fixed https://x.com/i/status/2075295876465979766


Got access to GPT 5.6 Sol Ultra, compared to Fable 5

468 points · 151 comments · r/OpenAI · by u/Accomplished_Whole_6

A user shares their experience comparing GPT-5.6 Sol Ultra to Anthropic's Fable 5, finding Sol Ultra to be a major upgrade from GPT-5.5 and notably more autonomous. When both models were asked to review a project and find errors in a fresh session, Sol Ultra found more issues, completed the audit faster, and demonstrated a better understanding of the full project. The user also noted Sol Ultra was 'ridiculously good at using the browser,' finding and fixing additional bugs unrelated to the original prompt.

Interesting Points
  • 5.6 Sol Ultra was tested in a fresh chat/session to review a project and find errors. It found more issues than Fable 5, did so faster, and had a better understanding of the full project.
  • It was described as 'ridiculously good at using the browser,' finding additional bugs unrelated to the original prompt which it then also fixed.
  • The 1.5x speed mode was noted as a significant quality-of-life improvement over Fable 5.
Top Comments

u/Dreki__ (270 points · permalink)

Even if Fable is slightly better on paper, Sol Ultra is going to win purely on usability. I am so tired of having to convince Claude that a SQL query deleting inactive user profiles is not an act of cyberterrorism.

u/Extra-Record7881 (50 points · permalink)

I am going to shift to gpt 5.6 as soon as my subscription ends, because I cant work with this level of constraints on fable 5, every message for my normal coding work, its getting flagged and I havw had enough of it, there is only so much i can do from my end to make sure I dont go back to gpt but it looks like dario is the toxic one in our relationship!

u/alew3 (48 points · permalink)

Just updated Codex and it now shows GPT 5.6

u/PsychMaster1 (58 points · permalink)

Well that took a lot less time that I thought it would.


I gave GPT-5.4, GPT-5.5, GPT-5.6 Sol, Terra and Luna the same 35-word Coca-Cola Zero brief

393 points · 51 comments · r/ChatGPT · by u/highsierraloft

A user ran a consistent frontend test across multiple ChatGPT models by giving each the same 35-word Coca-Cola Zero brief with reasoning cranked to the highest setting. The prompt asked for a beautiful landing page with at least five sections and a hero section, using only plain AI with custom design libraries allowed. Token budgets varied significantly across models, with Sol using over 200K tokens and Luna using under 100K.

Interesting Points
  • The exact prompt was: 'No skills are allowed. Create a beautiful landing page for Coca-Cola Zero using only plain AI. It can use custom design libraries. It must have at least five sections, with the hero section on top.'
  • Token budgets: Sol used 200,352 tokens, Terra used 154,574 tokens, Luna used 94,393 tokens.
  • Each model's output was hosted as a separate landing page for direct comparison, with an older Gemini 3.5 Flash included as a bonus comparison.
Top Comments

u/Ok_Researcher5981 (75 points · permalink)

Awesome, thanks for sharing! Seems like Gemini had the best looking images and multimodality (the interactive fizz sound)

u/smashbro64 (42 points · permalink)

Really cool to see the comparison. Thanks for spending your time and usage!

u/LakeRat (38 points · permalink)

Gemini killed it on this one. I'm impressed. I feel like even without taking the nano banana images into account it won here. Crazy that 3.5 Flash can hang with Sol on this.

u/DemonDookie (37 points · permalink)

As a graphic designer I agree, Gemini seems to have made the best page. It feels pretty much like an actual national brand product website.

The ChatGPT ones aren't bad but they would benefit a great deal from some directed refinement, especially with regards to product design/branding.


57 more Reddit stories

Updates: 06:00 AM PDT · 09:00 AM PDT · 12:00 PM PDT · 12:32 PM PDT · 03:00 PM PDT · 06:00 PM PDT