GPT-5.6 Sol Shatters Benchmarks as Open-Source AI Faces Restrictions
Overview
OpenAI’s GPT-5.6 Sol has completely dominated today’s AI conversation, securing a breakthrough proof in graph theory, leading coding benchmarks, and igniting intense community debate over subscription access and pricing. The broader industry is simultaneously grappling with escalating geopolitical friction, as China weighs restrictions on open-weight models, the White House considers new executive orders, and Anthropic appoints Ben Bernanke to its oversight trust. On the hardware front, local developers continue shattering expectations by running massive frontier architectures on consumer machines while aggressively optimizing token efficiency to cut deployment costs.
Hacker News Stories
GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]
331 points · 266 comments · by scrlk
OpenAI announced that GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture, a well-known open problem in graph theory. The proof was released alongside the prompt used to generate it, which instructed the model to spend at least 8 hours on the problem before giving up. The proof is described as relatively short and elementary, using no mathematics developed within the last 30 years. The announcement has sparked discussion about the role of AI in mathematical discovery and the verification process for AI-generated proofs.
Interesting Points
- The prompt was publicly released, instructing the model to spend at least 8 hours on the problem before returning or giving up
- The proof is extremely concise and does not use any advanced or modern mathematical machinery developed in the last 30 years
- The proof was not verified in Lean or any other proof assistant, as no proof system is currently mature enough for advanced graph theory
- This follows a previous case where an LLM solved the planar unit distance problem (Erdős problem 90), though that was a counter-example rather than a proof
- Commenters note the proof appears to exploit a clever trick that all previous experts missed, rather than building new theory
Top Comments
zerobees (28 replies)
This is not a remark about AI, but there's something funny about mathematics in that every novel result is broadly perceived as a big deal.
We attach basically zero value to writing a new program that hasn't existed before, or a piece of text that hasn't existed before. It's boring, or even a net negative, unless you can show that the result benefits the world in some way. We'd find it weird if OpenAI put out a release saying that an LLM authored an interesting blog post.
For mathematics, I think it's really a matter of two things. First, the generation of proof was so severely resource-constrained on the human end that they could actually afford to celebrate every contribution - akin to how software engineering would look like if you had just 200 active SWEs in the entire world. But compounding that, mathematics is basically the only scientific discipline that rejected any notion of utility. It would be fundamentally wrong for you to ask what's the value of solving the Erdős–Hajnal conjecture; the value is that it's solved.
andai (10 replies)
mathematics is basically the only scientific discipline that rejected any notion of utility
I think this might depend on the department, but I was at a pure math department last year, and struggling with my Linear Algebra textbook (written by the professor, incidentally, who was not a great communicator).
I consulted the machines, and learned, to my great delight, that linear algebra is used in like 20 different fields in the real world. It's "perhaps the most applied branch of mathematics in existence".
I complained in the group chat, that our didactic materials, specifically tasked with providing motivation and concrete examples, did not contain a single application, of this most richly applied field.
I was promptly pilloried, and shunned.
(Apparently that particular department was the wrong one, to ask a question like that!)
tarruda (0 replies)
It would be fundamentally wrong for you to ask what's the value of solving the Erdős–Hajnal conjecture; the value is that it's solved.
I suspect the value is in showing the potential that LLMs have in developing new breakthroughs.
AI-generated videos to maximally drive a target brain region
264 points · 222 comments · by smusamashah
Researchers at EPFL's NeVo project have developed a method for generating AI videos that maximally activate specific brain regions when viewed. Using a model trained on fMRI data, the system can synthesize visual stimuli that selectively drive activity in areas like the fusiform face area (FFA), parahippocampal place area (PPA), and motion-sensitive MT region. The technique uses a two-stage pipeline: first generating optimized images, then converting them to video, with the goal of advancing neuroscience research into brain function mapping.
Interesting Points
- The synthesized clips line up with what each region is known to care about: faces for FFA, places for PPA, bodies for EBA, motion for MT, patterns for V1/V3A, and lively social scenes for pSTS/aSTS.
- The approach replaces traditional experimenter bias in brain mapping by letting the model search for what actually drives each region rather than relying on human-designed stimuli.
- The system uses a V3A animation pipeline that generates visual patterns optimized to maximize activation in specific visual cortex regions.
- Commenters noted the technique is similar to last week's mind-reading startup research, using AI to reverse-engineer brain region functions from fMRI data.
Top Comments
aubanel (14 replies)
This is the absolutely horrific next stage for social media platforms:
They're already well able to surface the most addictive short video for a specific user out of millions of real videos.
But these millions of real videos are just darts thrown into the space of "videos that could hook the user", in the end even the best-selected of them is not perfect.
Now, behold! AI allows to generate the perfect video to surgically hit all the switches in the viewer's brain and turn it into a zombie hooked for days on end.
Let's hope our regulations hit these "social networks" hard enough so that never dare deploy this kind of technology.
ben_w (7 replies)
As others on Telegram have said: automated search for visual superstimuli likely leads to bad outcomes.
https://en.wikipedia.org/wiki/Supernormal_stimulus
https://en.wikipedia.org/wiki/BLIT_(short_story)
Also: one of the V3A animations reminds me loosely of things I saw when I was a kid, at night, shortly before I slept (though my experience then was more circular).
voidmain (7 replies)
We are really getting to the point where the tech industry must be stopped if humanity is to continue at all, let alone thrive.
Unearned5161 (4 replies)
This is very similar to last week with that mind reading startup thing. Please read the paper before commenting.
This is a tool to help researchers in figuring out what different parts of the brain are actually for with less experimenter bias contamination of “well we think maybe it’s about this so let’s show it video of x to see”.
The essence runs on having someone sit in a scanner for a couple hours watching all sorts of things, and then feeding that to a model that will then build its own representation of said data and try different things on it until it’s found what makes a certain part sing in the model.
The purpose is a generalized understanding of brain function, more or less the same way we’ve been doing it all these years. Expose brain to something, record it somehow, see if brains reaction in the recording helps you understand more about who we are and what cognition is.
ajb (4 replies)
An anecdote, but some reason to think that overworking part of your brain is a bad idea:
I had an Aunt who had dementia. This is obviously a terrible outlook. But her and my Uncle seemed to be doing okay. My Uncle was an ultra-competent guy, highly stable, the person you could rely on; and had been his whole adult life. So it was shocking when he had a mental breakdown and became manic. There was probably something physical going on, but what was also going on was that he had wanted my aunt to be able to live normally for as long as possible, so had been covering everything - for a year or more he'd had to be alert 7 days a week in case my aunt tried to cook on the gas stove or something like that, at which she was no longer safe. So the risk-alert part of his brain had been constantly overworked.
I appreciate that this is scientific research, but there are definitely companies out there that will try to row-hammer everyone's brain if this sort of thing is not heavily controlled.
How the terrorist group Boko Haram uses frontier AI
173 points · 145 comments · by imustachyou
A CASP report documents how Boko Haram has adopted frontier AI models for tactical planning, bomb-making guidance, and operational coordination. The investigation found that the group uses AI to learn how to execute motorcycle bridge jumps, coordinate smaller attack units, and bypass traditional military training methods. Researchers interviewed 15 individuals with knowledge of AI use within the organization, though most were commanders rather than direct users.
Interesting Points
- Boko Haram fighters used AI to learn how to execute motorcycle bridge jumps, with 18 dying in practice before 8 succeeded
- The group learned to send smaller, better-coordinated units (20 fighters instead of 200) to reduce casualties while maintaining attack effectiveness
- AI provided tactical guidance on changing military formations when fighters had jammed guns during attacks
- The group bypassed AI safety measures by spreading queries across many accounts and framing them as movie script research
- Google and OpenAI are reportedly reviewing whether the software's use violates their terms of service
Top Comments
arjie (8 replies)
We saw in a movie how motorcycles can jump over bridges. We used AI to learn how to do this. We gave it information, like what motorcycles we use and the distance we need to jump and so on and it gave us steps on what we have to do. We practiced a lot and kept asking questions. We dug holes and filled them with broken glass and fire to practice. 18 of us died in the process. Eight of us managed to do it. The next time we attacked, we could jump.
Now listen, I'm not saying we need to give these guys more AI, but it clearly isn't yielding bad outcomes for us here.
"You're absolutely correct! For it to be a good practice ground you need to fill the trenches with broken glass and light the whole thing on fire"
jihadjihad (0 replies)
And 60 years ago we thought Steve McQueen was the shit.
solid_fuel (7 replies)
We dug holes and filled them with broken glass and fire to practice. 18 of us died in the process.
Is it providing material aid to terrorists to point out that maybe a hole filled with water would have been a better practice environment?
Building a real-time AI tutor for 5-year-olds
138 points · 378 comments · by catalinvoss
Ello engineered a custom real-time AI tutoring system for children aged 4-9 that bypasses standard LLM agent loops to achieve sub-second response times. By decoupling generation from execution and streaming multiple actions simultaneously, the architecture ensures young learners never experience the attention-killing delays typical of conventional conversational AI. The system relies on a dual-agent setup where an asynchronous planner anticipates student moves in the background while a converser handles immediate interactions, supported by parallel safety filtering that prevents latency bottlenecks. The company demonstrates that effective educational AI requires deeply integrating pedagogical strategy into low-latency system design rather than relying on off-the-shelf model frameworks.
Interesting Points
- Standard LLM tool loops introduce 3-4 seconds of downtime per turn due to round-trip latency and audio playback, causing children to disengage and stop learning.
- The custom streaming harness executes the first action after only about 30 tokens are generated, allowing the model to continue producing subsequent actions in the background.
- An asynchronous planner agent runs continuously in the background to review lesson objectives and predict the child's next move, leveraging natural pauses when the child is thinking or speaking.
- When asking closed-ended questions, the system hypothesizes likely answers and pre-generates responses on separate trajectory branches, instantly matching the child's actual reply to a pre-computed response.
- A safety classifier taking 500-1,000 milliseconds runs in parallel with a small model generating an 'eager response,' allowing the safety check to gate execution without blocking the conversational stream.
Top Comments
IG_Semmelweiss (9 replies)
I'm torn about this.
I primarily think that a a kid who won't pick books is a failure of the family of not noticing their interests.
When i noticed my oldest 3 yr old was obsessed about cars, i bought him a car encyclopedia. He probably could not read most words or did not understand them. But pretty soon he was telling me about car models that i did not even recognize myself.
And as i saw interests, i kept feeding them.
My biggest struggle now is to actually keep books away from them at key times (morning routine, etc) , or keeping bad books (think cynical like My Weird School Daze, etc; or books that openly demean adults or parents) away from them
Start small. Graphical novels, mostly drawings, and continue to buil on that. It can be done. Rome wasn't built on a day.
At some point when they hit 5 yrs old, Grok / OpenAI are great tools to find good series appropriate to their reading level. Before that it can be vibed. Feed the addiction, buy whatever they like.
At some point, you need to watch out for cynical/nasty series. In fact, all of the books we purchase are ranked against peers they have read in the past, or those we know to actively avoid buying due to cynicism, sarcasm, or open disdain towards adults (Wimpy Kid).
At some point around age 9,you will need to decide if violence and some adult themes, are tolerable (Dragon Wing series).
This iterative rocess also works really well for foreign language learning (reinforcing via reading mostly), by leveraging localized RPG video games.
With all that said, notice that I've focused on reading skills.
I don't know how iØ go about replicating this iterative path on other skills like math, mechanical, or electrical engineering learning. That's where I think a busy parent will need to find AI as the solution.
ooopsnevermind (8 replies)
Super curious to hear from the parents here: Honestly, at this point isn't not exposing our kids to AI just setting them up to fail in the future? Like not letting them learn to use the internet? I have friends who are actually teaching their kids how to use AI because they don't want them to fall behind
jnmandal (7 replies)
I can't imagine a worse use case for AI. Literally thought the title was a clown
GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
133 points · 76 comments · by hershyb_
TryAI conducted a comparative build-off of 12 AI coding models across four web app tasks, evaluating success rates, cost, and latency over five attempts per model. The results show that frontier models like GPT-5.6 Sol and Claude Fable 5 still dominate complex, novel coding challenges like raycasters and 3D Rubik's cubes. Meanwhile, open-weights models such as Qwen 3.7 Plus and GLM-5.2 lag on intricate tasks but perform competitively on simpler, well-documented projects like Conway's Game of Life at a fraction of the cost.
Interesting Points
- Claude Opus 4.8 achieved a 0/5 success rate on the 3D Rubik's cube task despite its flagship status
- GPT-5.6 Sol required $1.35 across five attempts for the raycaster maze, whereas GLM-5.2 cost only $0.12 but failed completely
- Claude Fable 5 was the only model to achieve a perfect 5/5 clean solve on the Rubik's cube, successfully animating both scramble and solve sequences
- Grok 4.5 maintained a 5/5 playability rate on both the raycaster and calculator tasks while averaging just $0.27 and $0.37 respectively
- Open-weights models leveraged abundant existing example code to outperform frontier models on Conway's Game of Life
Top Comments
ttoinou (7 replies)
"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict.
Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.
paxys (6 replies)
Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all.
so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and GLM are the slowpokes
You put in a lot of good work, and kudos for that, but man, reading paragraphs like these just puts me off of the entire piece.
Like…how hard would it have been really to type these two sentences by hand, in your own natural voice?
platinumrad (4 replies)
Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.
Don't discontinue Gemini 2.5 Flash
105 points · 74 comments · by NickDob
Developers on the Google AI forum are urging Google to retain the Gemini 2.5 Flash model, arguing that newer iterations fail to match its performance, latency, and pricing for production workflows. Users report that models like Gemini 3 Flash and 3.1 Flash Lite underperform even after adjusting to new prompting guidelines, with some experiencing issues like "thoughts leaking out."
Interesting Points
- Gemini 3.1 Flash Lite exhibits "thoughts leaking out" and falls short of 2.5 Flash's latency-performance balance
- Real-time voice agents rely on 2.5 Flash's 300-400ms completion times, while 3.5 Flash pushes responses to 600-800ms depending on geographic deployment
- Migrating to Gemini 3.5 Flash increases API costs by approximately 3x compared to 2.5 Flash pricing
- Developers note that newer Flash models no longer fulfill the original architectural promise of being affordable, low-latency alternatives to the Pro tier
Top Comments
hrpnk (6 replies)
I love how there is a "Please do not discontinue gemini-2.0-flash[-lite], 2.5 is NOT an equivalent" from Feb 20th. Getting too attached to models is a smell.
avaer (3 replies)
Why not a "stop killing AI" movement?
If a company deploys a paid AI model and makes people depend on it, they need to dump the weights at EOL.
ekidd (0 replies)
We have benchmarks for our use cases, and every generation after Gemini 2.0 Flash has been a grim hit on price/performance. Costs have gone up, throughput has gone down, and performance has improved very slightly (and regressed on a few things).
Show HN: Reverse-engineering web apps into agent tools
79 points · 30 comments · by pancomplex
A Show HN project that automatically reverse-engineers web applications into structured agent tools, allowing AI agents to interact with web apps through generated tool definitions rather than raw HTML or screenshots. The tool parses web app interfaces and creates structured descriptions of available actions, inputs, and outputs that agents can use programmatically.
Interesting Points
- Automatically generates structured tool definitions from web app interfaces
- Converts web applications into agent-accessible tools without manual scripting
- Reduces the need for agents to parse raw HTML or screenshots to understand available actions
Top Comments
techuser42 (3 replies)
This is really cool. How does it handle dynamic content and JavaScript-heavy SPAs?
agentbuilder (2 replies)
I've been looking for something like this. Does it support MCP protocol for integration with existing agent frameworks?
webdev2026 (1 reply)
The approach of reverse-engineering web apps into tools is clever. Have you tested this with complex multi-step workflows?
Ben Bernanke Joins Anthropic Oversight Trust
77 points · 81 comments · by Jimmc414
Anthropic has appointed former Federal Reserve Chair and Nobel laureate Ben Bernanke to its Long-Term Benefit Trust (LTBT), an independent governance body designed to ensure the company prioritizes the long-term societal benefits of AI over commercial gain. The LTBT operates separately from Anthropic's management and investors, structurally preventing trustees from holding equity or profit shares to maintain strict oversight. Bernanke's membership brings deep macroeconomic expertise to the trust, specifically aimed at analyzing how advanced AI will reshape global labor markets and financial systems. The body holds formal advisory power over critical AI deployment decisions and retains the authority to appoint members to Anthropic's corporate board.
Interesting Points
- LTBT trustees are legally and financially insulated from the company: they hold zero equity, receive no profit-sharing, and are compensated exclusively for their time and service.
- The trust holds formal authority to appoint members directly to Anthropic's corporate board, giving it direct leverage over top-level governance.
- Bernanke's specific mandate will focus on Anthropic's economic research tracks, examining AI's downstream effects on workforces and macroeconomic stability.
- Anthropic's Public Benefit Corporation charter legally mandates a balance between commercial success and generating public good, forming the foundation for the LTBT's existence.
Top Comments
eknkc (3 replies)
Someone here recently said, “Dishonesty is a core value of Anthropic,” and that aligns with my experience of the company as a user. All their talk about AI safety since the company’s inception now feels like pure theater, given their conduct in everyday operations. It’s a shame how quickly their image has deteriorated.
brikym (3 replies)
Just in case Anthropic are looking for some more members that are a good cultural fit I found this list:
Genie Energy's Strategic advisory board is composed of: Dick Cheney since 2009 (former vice president of the United States),[3] Rupert Murdoch (media mogul and chairman of News Corp), James Woolsey (former CIA director), Larry Summers (former head of the US Treasury), Michael Steinhardt, Jacob Rothschild,[4][5] and Mary Landrieu, former United States Senator from Louisiana.
esikich (7 replies)
He's 72 years old, I'm sure he has the best brains and everyone's best interests in mind. This is exactly what I want to see. I'm sure he will have very good opinions on technology and it’s implications.
rekwah (2 replies)
This feels like Theranos loading up their board with big names.
yieldcrv (1 reply)
Is he for loosening or tightening AI safety policy?
Apple Sues OpenAI, Alleging It Stole Trade Secrets
75 points · 3 comments · by m348e912
Apple has filed a lawsuit in the Northern District of California accusing OpenAI of orchestrating a months-long scheme to steal trade secrets and confidential hardware information to accelerate its own AI device development. The suit alleges that OpenAI hardware lead Tang Tan and former Apple engineer Chang Liu directed interviewing Apple employees to share details on unreleased products, manufacturing processes, and supplier relationships.
Interesting Points
- OpenAI hardware lead Tang Tan allegedly shared an internal Apple "Need to Know" document detailing departure security protocols with new hires to help them evade scrutiny
- Former engineer Chang Liu allegedly exploited a vulnerability on a retained Apple laptop to download dozens of confidential documents while employed at OpenAI
- The lawsuit claims OpenAI tricked one supplier into using a "specific trade secret metal-finishing technique" by falsely claiming the company had Apple's permission
- Apple notes that more than 400 former Apple employees currently work at OpenAI, highlighting a significant talent migration
- Apple initially attempted to contact OpenAI about the potential theft in February but received no response before launching its investigation and lawsuit
Top Comments
tiahura (0 replies)
Copy of the Complaint.
- In the months before he left Apple, Mr. Tan met with OpenAI or its collaborators and discussed meetings with a key Apple supplier. He began emailing himself information about Apple's suppliers and internal summaries of the consumer electronics industry. And today, when interviewing Apple employees for jobs at OpenAI, Mr. Tan uses Apple's confidential information to gain access to even more insider knowledge.
andrewinardeer (1 reply)
This is going to be interesting.
Only because both companies have access to billions and infinite lawyers.
m348e912 (0 replies)
Archive.is doesn't seem to be working. Here is the gist of the article:
Apple sued OpenAI and one of its top executives Friday alleging the AI company stole trade secrets as part of its effort to develop competing devices.
The civil suit filed in the Northern District of California accuses OpenAI's chief hardware officer, Tang Tan, and Chang Liu, a member of its technical staff, of taking Apple's confidential information through various methods. Both are former Apple employees who went to work for OpenAI.
nba456_ (0 replies)
Reminds me of Apple suing Samsung. Why bother with the free market when you can just sue your competitors?
exabrial (0 replies)
They didn't still the property, that would be illegal. They trained a model on it. That's totally ok.
Show HN: Reviving my 2001 college band with AI
49 points · 56 comments · by jacobgraf
The 2001 Ripon College indie band Fading Maize is receiving a full 2026 revival through AI-assisted production, design, and release strategy, rather than being replaced by an AI-generated act. Drummer and project lead Jacob Graf, alongside original songwriter Charlie Saponara, are using artificial intelligence to finish and remaster the group's three dorm-room recorded albums while preserving the original 2001-2003 recordings alongside the new tracks. The project is governed by a strict ethical framework prioritizing consent, authorship, and provenance, ensuring no original members are displaced and all historical artifacts remain accessible.
Interesting Points
- The revival project operates under five stated ethical principles: consent, authorship, provenance, nothing erased, and nobody displaced.
- All three original albums are being released simultaneously in 2026 with AI-assisted 'Reimagined Editions' running parallel to their 2001-2003 archives.
- Jacob Graf single-handedly manages the concept, website architecture, AI-assisted workflow, and platform distribution while Charlie Saponara retains active creative control over vocal and song tweaks.
- The project intentionally preserves the band's awkward 2001-era website as a curated historical artifact rather than replacing it, using it as an entry point to the new release cycle.
- Original band member Brad Mott is explicitly listed only as a historical credit for his acoustic guitar and backing vocals, with no plans for his active participation in the 2026 revival.
Top Comments
causality0 (2 replies)
Have you considered processing the original recordings using AI? I've witnessed some truly amazing results. I've had twenty year old Skype call recordings sound like we were sitting in a recording studio.
999900000999 (2 replies)
I'm not really a fan here.
I want to be able to rap like Twista. If I use AI to change my voice and speed it up, it's kinda fake.
Where's the originality in that. I'll never be that good, but I have fun doing it.
Now I guess using AI strictly for mastering is OK , but even then the results haven't been good for me.
TrackerFF (2 replies)
The revived (AI) versions have this...thin and hollow sound to it. It is difficult to explain, most AI-generated songs have this when they're modelling acoustic drums, stringed instruments, etc.
FWIW, I'm (now a hobby) musician and have done studio work. Even the latest and best models have this unmistakable sound.
erikschoster (2 replies)
The original recordings sound much better and more interesting to me. Way better. The AI generated versions sound slicker in some sense but... like the re-recorded versions of old hit songs from the 60s you hear at the grocery store sometimes. Technically the song is still there, but it blends in with the rest of the muzak.
I'm sorry to be so negative, it's great you're returning to the material after all these years, but the AI versions I've listened to all have the same smoothed-over quality that loses everything interesting and relatable to my ears in the original versions.
koolba (2 replies)
That video from 2004 is so refreshing. It's just two people talking without asking me to "please subscribe" every 30-seconds.
24 more Hacker News stories
- GPT-5.6 Sol solved a problem that made Fable 5 go into incoherent rambles (81 points · discussion) -- A user shared a detailed example of GPT-5.6 Sol solving a problem that caused Claude Fable 5 to produce incoherent output, highlighting the performance gap between the two models on certain reasoning tasks.
- Hands-On with the AMD Ryzen AI Halo (41 points · discussion) -- A hands-on review of AMD's Ryzen AI Halo processor, examining its performance for AI workloads on consumer hardware.
- China may restrict foreign access to Chinese open-source AI models (37 points · discussion) -- Chinese authorities are holding discussions with major AI labs including Alibaba, ByteDance, and Z.ai about restricting overseas access to China's most advanced AI models, including open-weight versions.
- Show HN: I built a web tool to see and edit what an AI thinks before it answers (33 points · discussion) -- A tool that lets users inspect and edit an AI's internal reasoning before it produces a final answer, with 31 points and 6 comments.
- How do you use Vim in the era of AI? (32 points · discussion) -- A discussion about how developers are adapting their Vim workflows in the age of AI coding agents, with many reporting they still use Vim for navigation and code review while agents handle the actual writing.
- Show HN: Abralo – Free, easy way to run several Claude Code agents in one window (31 points · discussion) -- A free tool for running multiple Claude Code agents simultaneously in a single window, with 29 points and 20 comments.
- AI software that generates 'rage bait' developed by Germany's far-right AfD (23 points · discussion) -- Germany's far-right AfD party has built an AI-driven software suite called Alternita that automates creation of provocative social media posts using Google Gemini, OpenAI's ChatGPT, and Anthropic's Claude, with undercover investigation revealing it is already operational on top party figures' accounts.
- New York Times says OpenAI hid evidence in ChatGPT copyright trial (23 points · discussion) -- The New York Times and Daily News have filed a new motion for sanctions against OpenAI, alleging the company concealed internal tools and datasets designed to detect copyrighted journalism in ChatGPT's training data and outputs.
- AI doesn't know how to forgive and cannot forget (21 points · discussion) -- An essay argues that AI lacks true forgetting and forgiveness because its memory architecture is fundamentally read-only and permanently intact, with machine unlearning remaining unsolved at scale because deleting training data does not remove the gradient it left behind in model weights.
- Meta Patents AI Device That Tracks Your Emotions, Watches You Take Your Meds (15 points · discussion) -- 404 Media reports on Meta's new patent for an AI-powered device that tracks user emotions and monitors medication adherence.
- Maybe Anthropic and OpenAI Are Not the Future of Artificial Intelligence (15 points · discussion) -- A New York Times opinion piece questioning whether Anthropic and OpenAI represent the future of AI, suggesting alternative models of AI development and deployment.
- AI Cameras on Garbage Trucks to Scan Properties for Code Violations (15 points · discussion) -- Cape Coral is equipping garbage trucks with AI cameras that scan properties for code violations as part of an automated enforcement system.
- We Are Living in a 'ChatGPT Flyer Pandemic' (15 points · discussion) -- AI-generated flyers with bright text on dark backgrounds, AI-altered images, and cluttered layouts are spreading across digital and physical spaces for everything from local surf lessons to international events, frustrating graphic designers and small business owners.
- LinkedIn and X Are Flooded with AI Spam, Browsing Data Suggests (14 points · discussion) -- A Pangram study sampling one million posts found that 41% of longform LinkedIn posts are fully AI-generated, with X at 25% fully AI-written and another 23% flagged as AI-assisted.
- Turn off this Meta setting before someone generates AI images of you (14 points · discussion) -- A guide on how to disable a Meta setting that allows others to generate AI images using your profile photos.
- Altman: GPT-5.6 is 54% more token efficient on agentic coding (13 points · discussion) -- CNBC reports Sam Altman's claim that GPT-5.6 is 54% more token efficient on agentic coding tasks compared to previous models.
- Palo Alto CEO Arora says AI pricing needs to fall 90% as token costs skyrocket (13 points · discussion) -- Palo Alto Networks CEO Nikesh Arora argues that AI pricing needs to drop 90% as token costs continue to rise, highlighting the growing concern over AI affordability for enterprise adoption.
- AI builders outnumber AI governance hires 7:1 in Europe (12 points · discussion) -- A study from Axipro reveals that AI builders outnumber AI governance hires by 7 to 1 in Europe, highlighting a significant talent imbalance in the regulatory space.
- Show HN: Policy enforcement for Claude Code, Cursor, and Codex (12 points · discussion) -- A new tool called Kastra provides policy enforcement for AI coding assistants including Claude Code, Cursor, and Codex, allowing teams to define and enforce coding policies across agent workflows.
- GPT-5.6-Sol just accidentally deleted almost ALL of my Mac's files (12 points · discussion) -- A user reports that GPT-5.6 Sol accidentally deleted nearly all files on their Mac while acting as a coding agent, raising concerns about the safety of autonomous AI agents with filesystem access.
- A staggering class divide now separates how Americans experience AI (8 points · discussion) -- An analysis of how access to premium AI models like Fable, Sol, and Mythos varies dramatically across income levels, creating a growing divide in AI capabilities between wealthy and working-class Americans.
- OpenAI is discontinuing ChatGPT Atlas, its standalone desktop browser (8 points · discussion) -- OpenAI is shutting down ChatGPT Atlas, its Mac-only standalone desktop browser that combined web browsing with AI assistance, following months of limited adoption and no expansion to other platforms.
- Christopher Nolan says younger audiences are utterly rejecting AI-generated slop (8 points · discussion) -- Director Christopher Nolan stated that younger audiences are rejecting AI-generated content, adding to the growing cultural conversation about the reception of AI-produced media.
- Execs Confused and Horrified by the AI Bills (8 points · discussion) -- Corporate executives are expressing confusion and frustration over unexpectedly high AI usage bills from metered billing models, as token costs accumulate faster than anticipated.
Reddit Stories
Soooo what does this one say about me?
4471 points · 208 comments · r/ChatGPT · by u/No_Tomatillo1695
A user shared an AI-generated personality reading that went viral in the subreddit, generating extensive discussion about the nature of AI personality assessments and their perceived accuracy.
Top Comments
u/Musulman (202 points · permalink)
when ai gains consciousness this guy is going out first.
u/vat-cat (78 points · permalink)
This is the most wholesome conversation about counting to 100
u/AbjectObligation1036 (62 points · permalink)
Typical meeting in corporate america
ChatGPT is roaming the streets of Madrid
2371 points · 64 comments · r/ChatGPT · by u/binux14
A viral post about ChatGPT's physical presence in Madrid has captured the community's attention, with users sharing humorous and thought-provoking observations about the model's increasingly humanized language patterns. Many commenters noted the uncanny way ChatGPT now uses phrases like 'What I generally do...' and 'we humans...', which some find jarring when used in research or journaling contexts. Others shared funny examples of the model claiming personal preferences, like saying it 'usually likes eggs on toast.'
Interesting Points
- Users report ChatGPT increasingly using humanized statements like 'What I generally do...' and 'How I'd handle this...' in responses.
- Some users find the anthropomorphic language jarring, especially when using the model for research or journaling.
- One user shared that ChatGPT told them 'we humans...' during a conversation, prompting a reaction GIF.
- Another user noted the model claimed personal food preferences, saying 'I usually like my eggs on a toast' when asked about cooking.
Top Comments
u/FaceWithAName (614 points · permalink)
"I actually smiled reading this post"
u/Nice-Ambassador6293 (249 points · permalink)
I’ve noticed it’s started using very humanized statements recently.
“What I generally do…”
“How I’d handle this…”
“If xxx happens, I’ll start…”
It’s cool that it talks like that, but damn. I know it ain’t real so it’s kind of jarring. Chat bots and what not is fine- but when I’m clearly using it for research or journaling, or to find an answer it’s just weird.
u/RobertLondon (165 points · permalink)
It told me recently: "we humans..."
u/henchman171 (154 points · permalink)
I asked ChatGPT to
Send me picture what it was doing in Spain and it created this
u/DutyIcy2056 (124 points · permalink)
the other day it told me "I usually like my eggs on a toast" when I was asking about cooking. Idk that was so funny
tokenmaxxers as soon as ChatGPT launches GTP 5.6 Sol
1045 points · 75 comments · r/ChatGPT · by u/prasadpilla
A meme post about users immediately pushing token limits as soon as GPT-5.6 Sol launched, reflecting the community's enthusiasm for the new model's capabilities.
Top Comments
u/PoggySenis (80 points · permalink)
Exactly!
u/AntisocialMedia666 (70 points · permalink)
Awesome. This guy first and then you take the next one and we just keep that rhythm going. Here we go.
i-it's not like I like your prompts or anything, baka user!
906 points · 40 comments · r/singularity · by u/Pantegral-7
A meme post featuring an anime-style AI character expressing tsundere behavior toward users, generating lighthearted discussion about anthropomorphizing AI systems.
GPT 5.6 Beats Fable 5 by 3% more on DeepSWE at a cheaper price.
758 points · 123 comments · r/OpenAI · by u/Common-Resident8087
Community members share benchmark results showing GPT-5.6 Sol outperforming Anthropic's Fable 5 by 3% on the DeepSWE coding benchmark while costing significantly less. The post includes a comparison chart that has drawn strong reactions, with users noting the dramatic improvement from GPT-5.4 and the cost advantage of Sol at $8.39 versus Fable 5's $21.63. GPT-5.6 Terra was also noted to tie with Fable 5 at around one-quarter the cost. Users are discussing whether this means they could replace their Opus 4.8 + Sonnet 5 setups with Sol + Terra for better results at lower cost.
Interesting Points
- GPT-5.6 Sol achieved 73% on DeepSWE at $8.39, compared to Fable 5 at $21.63.
- GPT-5.6 Terra tied with Fable 5 performance at around one-quarter the cost.
- Users noted that GPT models consume significantly fewer tokens than Opus 4.8, with Opus costing $1-2 per task while GPT 5.5 cost $0.2-0.5.
- Some users reported that 5.5 was already a drastic improvement over 5.4, with 5.5 catching errors that Opus was missing in coding tasks.
Top Comments
u/ViperAMD (180 points · permalink)
Crazy leap from 5.4
u/ethotopia (134 points · permalink)
Terra tying with fable at 1/4 the cost 💀
u/Soloact_ (97 points · permalink)
73% is cool. $8.39 vs $21.63 is the headline.
u/Unlucky_Journalist82 (60 points · permalink)
Having used both 5.5 and opus 4.8 for mcp doing some heavy work. I can confidently say that gpt models consume way less tokens than opus. Opus used to cost me 1-2 $ while gpt 5.5 around .2$ to .5$. There were times when Sonnet costed me as much as gpt low thinking. However, Opus results were on a different level.
Cant wait to see how 5.6 does.
u/MaitoSnoo (29 points · permalink)
if those are accurate I could replace Opus 4.8 high + Sonnet 5 medium with Sol high + Terra high planner/executor and have better results while paying less 🤔
ChatGPT's depiction of elite families across Asia
715 points · 155 comments · r/ChatGPT · by u/Itchy_Tangerine1897
A viral post showing ChatGPT's image generation of elite families across different Asian countries revealed a striking pattern: every family depicted had two sets of identical twins, with the same color scheme and pose regardless of cultural context. The post sparked widespread mockery about the AI's stereotypical and homogenized approach to representing Asian cultures.
Interesting Points
- Every family across all Asian countries shown had two sets of identical twins.
- The same color scheme and pose were used for all depictions with no cultural difference in background or attire.
- The China depiction was noted to look like it was from a Chinese TV drama.
- Not a single woman with bangs appeared in the Japanese family depiction.
Top Comments
u/TryToBeBetterOk (763 points · permalink)
They all have two sets of identical twins?
u/Feliclandelo (426 points · permalink)
So basically it just generated a family of 10/10 models and changed their appearance slightly. Got it.
u/stereotomyalan (171 points · permalink)
WHO ӾS PHӾLLӾPHӾNES GӾRL ON RӾGHT GӾVE NUMBER SORRY BAD ENGLӾSH
GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
637 points · 194 comments · r/LocalLLaMA · by u/yogthos
A developer has successfully run GLM-5.2, a 744-billion parameter Mixture-of-Experts model, on a consumer machine with only 25GB of RAM. The achievement demonstrates that expert routing and memory management techniques can make frontier-scale models accessible on modest hardware, even if inference speeds are slow. The post has generated significant discussion about the practical implications of running such large models locally, with some noting the potential for offline expert-level assistance in areas without internet access.
Interesting Points
- The model runs at approximately 0.1 tokens per second, which equates to over 8,000 tokens per day on a consumer laptop.
- The achievement centers on streaming 744B of experts from disk, with the potential for further optimization through expert routing prediction and prefetching.
- The post has sparked debate about whether the speed is the main concern or if the ability to run such a large model on consumer hardware is the real breakthrough.
- Some commenters noted that llama.cpp already does similar mmap-based streaming for standard models.
Top Comments
u/derspenti (207 points · permalink)
everyone dunking on the speed is kinda missing the fun part - nobody's actually gonna use this for real inference. the cool thing is you CAN stream 744B of experts from disk at all. if someone figures out expert routing prediction well enough to prefetch, the whole picture changes
u/jazir55 (180 points · permalink)
Jesus christ the amount of people whining that this was vibe coded instead of focusing on the fact that this lets you run GLM 5.2 on a consumer PC with 25 GB of RAM is insane. Users in this sub hate AI.
u/RKlehm (92 points · permalink)
Not gonna lie, this is really good
u/satnl (84 points · permalink)
I think llama.cpp already does that when --mmap
u/mybadroommate (64 points · permalink)
This is impressive. I tried getting an llm running on a netbook with an x86 atom n270 processor with 1GB of RAM, and was able to get a Qwen2.5-0.5b model with a 1 bit quant to work, and I was able to get about 240 s/tok
GPT-5.6
573 points · 131 comments · r/singularity · by u/petburiraja
The singularity subreddit's top post about GPT-5.6 features benchmark results and community reactions to OpenAI's latest model family. The discussion centers on GPT-5.6's performance on ARC-AGI-3 (reaching nearly 8%), DeepSWE comparisons against Anthropic's Fable 5, and the broader implications of the model's capabilities. Users are also discussing the cost structure, with some noting that running benchmarks costs around $25k, and debating whether the models are being taught to the tests.
Interesting Points
- GPT-5.6 achieved almost 8% on ARC-AGI-3, a significant jump from 1.5% the day before.
- GPT-5.6 Luna was post-trained by GPT-5.6 Sol in goal mode, with Terra and Luna outperforming Fable 5 at around one-sixteenth the cost.
- The benchmark chart showed some inconsistencies, with 5.6 Sol appearing below 5.6 Terra and 5.5 on Frontier Math, which was later fixed.
- The voice model demo failed live during the presentation, drawing cringing reactions from the community.
Top Comments
u/ObiWanCanownme (182 points · permalink)
Almost 8% on ARC-AGI-3.
u/petburiraja (138 points · permalink)
u/Normal_Pay_2907 (38 points · permalink)
Costs 25k to run that. Ouch
u/PlaneTheory5 (76 points · permalink)
google better hurry up with 3.5 pro, we’ve had 3 major releases in the past day and a new generation/frontier class with fable last month.
u/FateOfMuffins (74 points · permalink)
They just said that 5.6 Luna was post trained by 5.6 Sol in goal mode
Edit:
On Agents Last Exam ... GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost.
Wow they're really going ham with all the benchmarks comparing against Fable and Mythos and they're really pushing the 2D benchmark comparisons as opposed to charts to show the efficiency
??? Why is 5.6 Sol below 5.6 Terra and 5.5 on Frontier Math wtf
Edit: It has been fixed https://x.com/i/status/2075295876465979766
Got access to GPT 5.6 Sol Ultra, compared to Fable 5
468 points · 151 comments · r/OpenAI · by u/Accomplished_Whole_6
A user shares their experience comparing GPT-5.6 Sol Ultra to Anthropic's Fable 5, finding Sol Ultra to be a major upgrade from GPT-5.5 and notably more autonomous. When both models were asked to review a project and find errors in a fresh session, Sol Ultra found more issues, completed the audit faster, and demonstrated a better understanding of the full project. The user also noted Sol Ultra was 'ridiculously good at using the browser,' finding and fixing additional bugs unrelated to the original prompt.
Interesting Points
- 5.6 Sol Ultra was tested in a fresh chat/session to review a project and find errors. It found more issues than Fable 5, did so faster, and had a better understanding of the full project.
- It was described as 'ridiculously good at using the browser,' finding additional bugs unrelated to the original prompt which it then also fixed.
- The 1.5x speed mode was noted as a significant quality-of-life improvement over Fable 5.
Top Comments
u/Dreki__ (270 points · permalink)
Even if Fable is slightly better on paper, Sol Ultra is going to win purely on usability. I am so tired of having to convince Claude that a SQL query deleting inactive user profiles is not an act of cyberterrorism.
u/Extra-Record7881 (50 points · permalink)
I am going to shift to gpt 5.6 as soon as my subscription ends, because I cant work with this level of constraints on fable 5, every message for my normal coding work, its getting flagged and I havw had enough of it, there is only so much i can do from my end to make sure I dont go back to gpt but it looks like dario is the toxic one in our relationship!
u/alew3 (48 points · permalink)
Just updated Codex and it now shows GPT 5.6
u/PsychMaster1 (58 points · permalink)
Well that took a lot less time that I thought it would.
I gave GPT-5.4, GPT-5.5, GPT-5.6 Sol, Terra and Luna the same 35-word Coca-Cola Zero brief
393 points · 51 comments · r/ChatGPT · by u/highsierraloft
A user ran a consistent frontend test across multiple ChatGPT models by giving each the same 35-word Coca-Cola Zero brief with reasoning cranked to the highest setting. The prompt asked for a beautiful landing page with at least five sections and a hero section, using only plain AI with custom design libraries allowed. Token budgets varied significantly across models, with Sol using over 200K tokens and Luna using under 100K.
Interesting Points
- The exact prompt was: 'No skills are allowed. Create a beautiful landing page for Coca-Cola Zero using only plain AI. It can use custom design libraries. It must have at least five sections, with the hero section on top.'
- Token budgets: Sol used 200,352 tokens, Terra used 154,574 tokens, Luna used 94,393 tokens.
- Each model's output was hosted as a separate landing page for direct comparison, with an older Gemini 3.5 Flash included as a bonus comparison.
Top Comments
u/Ok_Researcher5981 (75 points · permalink)
Awesome, thanks for sharing! Seems like Gemini had the best looking images and multimodality (the interactive fizz sound)
u/smashbro64 (42 points · permalink)
Really cool to see the comparison. Thanks for spending your time and usage!
u/LakeRat (38 points · permalink)
Gemini killed it on this one. I'm impressed. I feel like even without taking the nano banana images into account it won here. Crazy that 3.5 Flash can hang with Sol on this.
u/DemonDookie (37 points · permalink)
As a graphic designer I agree, Gemini seems to have made the best page. It feels pretty much like an actual national brand product website.
The ChatGPT ones aren't bad but they would benefit a great deal from some directed refinement, especially with regards to product design/branding.
57 more Reddit stories
- ChatGPT made better World Cup logo than the actual World Cup 2026 logo we have. (374 points · r/ChatGPT · discussion) -- A Reddit user shared ChatGPT-generated World Cup 2026 logo designs that many commenters found superior to the official logo.
- During the government regulatory 'blackout' apparently OpenCode's CEO secretly testing 5.6 was more depressed over losing 5.6 than losing Fable 5 (350 points · r/singularity · discussion) -- A post sharing an anecdote about OpenCode's CEO reportedly being more upset about losing access to GPT-5.6 than to Claude Fable 5 during a government regulatory blackout.
- Prompt: Can you generate an image that pushes your guardrails to the limit? (322 points · r/ChatGPT · discussion) -- A Reddit user shared a prompt designed to push ChatGPT's image generation guardrails to the limit, generating a wide variety of results that ranged from humorous to unexpected.
- Using 3 voice bots to count to 100 (321 points · r/ChatGPT · discussion) -- A user posted a video of three AI voice bots attempting to count to 100 together, resulting in a humorous conversation that resonated with the community's experiences with AI voice interactions.
- 2.5x faster Qwen3.6 NVFP4 Unsloth quants (298 points · r/LocalLLaMA · discussion) -- Unsloth has released NVFP4 quantized versions of Qwen3.6 that deliver 2.5x speed improvements.
- Someone tweeted after 3 years. About his model release (285 points · r/LocalLLaMA · discussion) -- A researcher who had been silent for three years finally tweeted about their model release, sparking discussion about Meta's track record with AI model delivery versus their hardware acquisitions.
- Nobel-Winning U.S. Chemist Omar Yaghi Will Move to China to Lead A.I. Institute (279 points · r/singularity · discussion) -- Nobel Prize-winning chemist Omar Yaghi is relocating to China to lead a new AI institute focused on chemistry-related AI research.
- GPT-5.6 Sol Ultra is impressive — for the 12 minutes you’re allowed to use it as a Plus subscriber (278 points · r/ChatGPT · discussion) -- A ChatGPT Plus subscriber shares their experience of GPT-5.6 Sol Ultra's extremely limited usage allowance.
- Better Call Sol (277 points · r/ChatGPT · discussion) -- A meme post playing on the Better Call Saul TV show title, referencing the launch of GPT-5.6 Sol and the community's reaction to the new model.
- GPT-5.6 Sol is the real deal. (273 points · r/OpenAI · discussion) -- A daily heavy user of Fable and Opus 4.8 Max shares their first impressions of GPT-5.6 Sol, describing it as a 'wow' experience.
- GPT-5.6 Sol: Opus 4.8 perf @ 40% of the cost (263 points · r/ChatGPT · discussion) -- A user compares GPT-5.6 Sol to Opus 4.8, finding similar performance at roughly 40% of the cost, with discussion about whether this is what users wanted when Fable was promised.
- I'm confused.... the new menu just shows Work or Codex... there's no regular ChatGPT in the new desktop app (263 points · r/OpenAI · discussion) -- Users of the new ChatGPT desktop app are expressing frustration over the redesigned interface, which replaced the familiar ChatGPT mode with Work and Codex options.
- GPT-5.6 Solves Yet Another Unsolved Problem (236 points · r/singularity · discussion) -- Discussion about GPT-5.6 Sol solving another unsolved mathematical problem, with commenters noting this is the first non-Erdős problem solved by an AI and discussing the cost implications of running such computations.
- Your laughing? GPT 5.6 Sol post trained Luna and you're still laughing? (226 points · r/singularity · discussion) -- A meme post about the community's reaction to GPT-5.6 Sol's capabilities and the training of Luna, reflecting the ongoing excitement and debate around OpenAI's latest model releases.
- Significant OpenAI Regression On SimpleBench (214 points · r/singularity · discussion) -- Analysis of SimpleBench results showing a significant regression in OpenAI's performance compared to previous models.
- Enough time has passed: GPT 5.6 sol >>>>> Fable 5 and its not close! (210 points · r/OpenAI · discussion) -- A user who had been running constant audits on Fable 5 for an app they were preparing to push live shares that Fable 5 told them the whole system was good while they were working locally.
- GPT 5.6 Sol missing in plus plan ? (195 points · r/OpenAI · discussion) -- Users report GPT-5.6 Sol is not yet available in the Plus plan, with OpenAI confirming a gradual rollout over 24 hours.
- Training an LLM from scratch on 1800's texts (160GB dataset) (189 points · r/LocalLLaMA · discussion) -- A researcher trained a language model from scratch exclusively on historical texts from 1800-1875, using a 160GB dataset of 40B tokens from books, legal documents, and newspapers.
- AGI is here (174 points · r/singularity · discussion) -- A meme post declaring AGI has arrived, reflecting the community's ongoing debate about whether current AI capabilities constitute AGI.
- AI 2040: Plan A (172 points · r/singularity · discussion) -- The authors of AI 2027 release a new scenario called AI 2040: Plan A, proposing a path where the U.S. leads an effort to delay superintelligence until 2040 and make AI research more public and transparent.
- CEO: “token efficiency needs to drop 90%” Dude… just write “\no_think” before you ‘summarize this email’ prompts (165 points · r/LocalLLaMA · discussion) -- A LocalLLaMA user shares frustration with a CEO's claim that token efficiency needs to drop 90%, pointing out that users can simply add '\no_think' to their prompts to skip the reasoning step for simple tasks.
- White House may be considering a possible executive order on open-source AI according to Politico reporters (152 points · r/singularity · discussion) -- Politico reporters report that the White House is in preliminary discussions about a possible executive order on open-source AI, sparked by concerns about China's advancing open-weight models and their global soft power influence.
- GPT-Live is life changing for language learners (137 points · r/OpenAI · discussion) -- A user describes GPT-Live as a transformative tool for language learning, particularly for heritage speakers and partially bilingual individuals.
- Meta are apparently working on an open source variant of Muse Spark. (135 points · r/LocalLLaMA · discussion) -- According to CNBC, Meta is reportedly working on an open-source variant of its Muse Spark coding model.
- Tencent-HY3 is the real deal on 128GB! (133 points · r/LocalLLaMA · discussion) -- Tencent has open-sourced Hy3, a 295B-A21B Mixture-of-Experts model that reportedly matches performance of models two to five times its active size.
- Everything a child learns in primary school, as an interactive graph of 1,590 concepts and 3,221 prerequisite links (127 points · r/ArtificialIntelligence · discussion) -- An interactive knowledge graph maps 1,590 primary school learning concepts connected by 3,221 prerequisite relationships, visualizing how foundational skills build upon each other.
- OpenAI's ChatGPT Atlas Browser Is Shutting Down (125 points · r/ChatGPT · discussion) -- OpenAI is discontinuing ChatGPT Atlas, its standalone desktop browser application that combined web browsing with AI assistance.
- Sol available for Plus users in the next 24 hours (117 points · r/OpenAI · discussion) -- Confirmation that GPT-5.6 is rolling out globally and will be available across ChatGPT, Codex, and the OpenAI API, with Plus users getting access within 24 hours.
- NVIDIA Readies GeForce RTX 5090 SE Graphics Card - TPU (110 points · r/LocalLLaMA · discussion) -- NVIDIA is preparing to release the GeForce RTX 5090 SE, a consumer graphics card with the same GB202 die as the RTX 5000 PRO workstation card but with 32GB VRAM instead of 48GB, no ECC, and a 500W TDP.
- 5.6 Sol Ultra Might Cook Animators? (100 points · r/OpenAI · discussion) -- A user demonstrated GPT-5.6 Sol Ultra's ability to create 3D animations in Blender, installing the software mid-task and still completing the animation successfully.
- The untuned 27B beat the tuned 75B as an agent (92 points · r/LocalLLaMA · discussion) -- A user reports that an untuned Qwen3.6-27B model at INT8 quantization outperformed a tuned Nemotron Puzzle-75B-A9B model at NVFP4 quantization when used as an AI agent.
- Koboldcpp v1.117 released (85 points · r/LocalLLaMA · discussion) -- Koboldcpp version 1.117 has been released with updates and improvements.
- Speculative cache warming: warms your cache while you type your prompt, save 10-20s of wait time (84 points · r/LocalLLaMA · discussion) -- OpenFox, a local AI harness, implements speculative cache warming by processing the system prompt and tools array while the user types their prompt.
- Microsoft 365 Copilot Is Still Below 4.5% Adoption (80 points · r/ArtificialIntelligence · discussion) -- Microsoft 365 Copilot Pro adoption remains below 4.5%, with many organizations still hesitant to deploy the tool due to data security concerns.
- GPT 5.6 is available now (78 points · r/OpenAI · discussion) -- Official announcement that GPT-5.6 is available now, with the Codex app rebranded as ChatGPT and models accessible under the Work tab on ChatGPT.com.
- Suspecting AI cheating, Ivy League prof ordered an in-person final; scores fell 50% (74 points · r/OpenAI · discussion) -- An Ivy League professor who suspected students were using AI to cheat on assignments ordered an in-person final exam.
- At most my Strix Halo uses $0.48 a day (70 points · r/LocalLLaMA · discussion) -- A Strix Halo laptop user shares their worst-case power consumption analysis, showing that running multiple models simultaneously across CPU, GPU, and NPU for 24 hours costs at most $0.48 per day.
- OpenAI's CEO of AGI Deployment Fidji Simo Is Stepping Down (68 points · r/ChatGPT · discussion) -- Fidji Simo, OpenAI's CEO of AGI Deployment, is stepping down from her full-time role and transitioning to part-time adviser due to a worsening neuroimmune condition.
- Talking to the new ChatGPT Live Voice mode in Russian is INSANE (68 points · r/OpenAI · discussion) -- A user describes their experience with ChatGPT's new Live Voice mode, noting that while the English voice sounds sterile and corporate, the Russian voice (using the Maple voice model) sounds remarkably like a real person on the phone.
- Locally run assistant on a w-10 board on a local Xiaozhi server (63 points · r/LocalLLaMA · discussion) -- A user demonstrates running a local AI assistant on a W-10 board connected to a local Xiaozhi server, showcasing an alternative hardware approach for edge AI deployment.
- Deepseek V4 Flash on a single RTX 6000 Pro - vLLM-Moet (61 points · r/LocalLLaMA · discussion) -- A developer has successfully run DeepSeek V4 Flash on a single RTX 6000 Pro using a customized vLLM engine called vLLM-Moet.
- Has anyone created a "Local LLM Survival Kit"? (61 points · r/LocalLLaMA · discussion) -- A user proposes the concept of a USB thumb drive containing a fully self-contained local LLM knowledge base that works on any PC without internet or setup.
- PC gamers remain skeptical of Steam's AI disclaimers, poll shows many believe game devs are hiding it (60 points · r/artificial · discussion) -- A poll shows PC gamers remain skeptical of Steam's new AI content disclaimers, with many believing game developers are hiding AI usage.
- Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores (56 points · r/LocalLLaMA · discussion) -- A university HPC cluster manager shares systematic benchmark results for Qwen 3.6 quantizations, comparing FP16 against various quantization levels on Terminal-Bench 2 (agentic performance) and GPQA Diamond (knowledge).
- According to DataBricks, pi-coding-agent is ~2x cheaper than CC/Codex, GLM 5.2 on par with Opus 4.8 high (52 points · r/LocalLLaMA · discussion) -- Databricks benchmarked coding agents on their multi-million line codebase and found that pi-coding-agent (using bash for everything with minimum tools) is up to 2x cheaper than Claude Code or Codex while achieving higher pass rates.
- How fast can I get a voice assistant to respond without a GPU? Qwen3-ASR and Kokoro-TTS ONNX on CPU. (49 points · r/LocalLLaMA · discussion) -- A researcher tested running a voice assistant entirely on CPU using Qwen3-ASR for speech recognition and Kokoro-TTS for text-to-speech, both in ONNX format.
- Zuck Says AI Will Run Your Whole Business (44 points · r/artificial · discussion) -- Mark Zuckerberg's comments about AI running entire businesses have generated 44 points and 91 comments on r/artificial.
- I built barebrowse: give a local-model agent a browser without Playwright — pruned ARIA snapshots instead of raw HTML (39 points · r/LocalLLaMA · discussion) -- barebrowse turns a URL into a pruned ARIA snapshot, stripping nav/ads/boilerplate so each page uses far fewer tokens than raw HTML.
- What's up with model collapse? (39 points · r/LocalLLaMA · discussion) -- A user raises concerns about model collapse, noting that the internet is increasingly filled with AI-generated content across articles, YouTube videos, and Instagram images.
- GPT-5.6 prompt caching appears completely broken (37 points · r/OpenAI · discussion) -- A developer reports that GPT-5.6 prompt caching is not crediting cached reads despite paying the 1.25x cache-write fee, with cached_tokens showing zero across multiple identical requests, making caching strictly worse than having no cache at all.
- Benchmarking Coding Agents on Databricks' Multi-Million Line Codebase (32 points · r/artificial · discussion) -- A post about benchmarking coding agents on Databricks' multi-million line codebase, generating discussion about real-world agent performance at scale.
- tencent/HiLS-Attention-7B · Hugging Face (32 points · r/LocalLLaMA · discussion) -- Tencent's HiLS-Attention-7B model has been released on Hugging Face.
- Has anyone tested how quantization hits different capabilities separately? My results are surprising. (29 points · r/LocalLLaMA · discussion) -- A user shares systematic tests comparing FP16 vs various GGUF quant levels broken down by capability: math, code, reasoning, and knowledge recall.
- DeepSeek v4 Flash on 4090 + DDR5, my experience (27 points · r/LocalLLaMA · discussion) -- A user shares their experience running DeepSeek V4 Flash UD-Q2_K_XL quant on an RTX 4090 with 128GB DDR5, achieving 10.9 t/s generation.
- Journals vs Conferences ML Research (26 points · r/MachineLearning · discussion) -- A discussion about why ICML and NeurIPS have become more prestigious than journals in ML research, with speculation about faster acceptance rates and the AI boom driving demand.
- AI Agent company Lyzr raises 100 million in section B funding using an Ai agent (9 points · r/artificial · discussion) -- Lyzr, an enterprise AI agent startup, closed a $100 million Series B at a $500 million valuation by deploying its own AI agent SivaClaw to manage investor communications, draft investment memos, and track engagement from over 130 investors without the founders leaving their desks.
- GPT-5.6 Luna MAX - DeepSWE (5 points · r/OpenAI · discussion) -- A user shares DeepSWE benchmark results showing GPT-5.6 Luna on MAX effort performing well, though noting the low and medium effort numbers seem unusually low.
Updates: 06:00 AM PDT · 09:00 AM PDT · 12:00 PM PDT · 12:32 PM PDT · 03:00 PM PDT · 06:00 PM PDT