· 05:30 PM PDT

Frontier Silicon and Massive Models Reshape the AI Landscape

Overview

Apple’s latest M-series silicon and OpenAI’s custom Jalapeño ASIC underscore a fierce hardware race focused on boosting on-device and inference efficiency. Frontier model scale continues accelerating, highlighted by Qwen’s new Flash variant, rumored OpenAI pretraining runs surpassing 10 trillion parameters, and Anthropic’s staggering $30 trillion market projection. Beyond the lab, practical shifts are reshaping daily use and enterprise strategy, as OpenAI reinstates usage caps, new studies highlight AI’s disproportionate impact on entry-level jobs, and mounting concerns over trust and ROI push companies to prioritize measurable returns. Meanwhile, the open-source community keeps pace with breakthroughs in local inference, persistent agent harnesses, and advanced quantization that make edge AI increasingly viable.


Hacker News Stories

Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute

925 points · 880 comments · by interpol_p

Apple M6 and M5 Ultra chips

Apple has introduced the M6 and M5 Ultra silicon chips, marking significant advancements in performance and on-device AI capabilities. The M6, Apple's first 2-nanometer chip, debuts in the new Mac mini with a 12-core CPU and a Dual 16-core Neural Engine designed to accelerate everyday tasks and local AI workflows. Meanwhile, the M5 Ultra powers the new Mac Studio using a novel quad-die architecture to deliver massive unified memory bandwidth and unprecedented GPU compute for professional and frontier AI applications. Together, these chips enable developers and users to run and fine-tune large language models entirely on-device with industry-leading power efficiency.

Interesting Points
  • M6's new CPU complex splits into 2 super cores, 4 performance cores, and 6 efficiency cores, delivering up to 1.2x faster multithreaded performance than the M5.
  • The M5 Ultra utilizes a first-of-its-kind quad-die design that connects two dual-die M5 Max chips via UltraFusion, achieving an inter-die bandwidth exceeding 4.4TB/s.
  • M5 Ultra supports up to 512GB of unified memory with 1.2TB/s of bandwidth, enabling the local execution of LLMs with hundreds of billions of parameters without cloud dependency.
  • Both chips integrate Neural Accelerators directly into their GPU cores, boosting peak AI compute by nearly 30% on the M6 and up to 4.5x on the M5 Ultra compared to their predecessors.
  • The M5 Ultra's media engine includes hardware-accelerated AV1 decode and four dedicated ProRes encode/decode engines, while Apple Intelligence features are slated for release with macOS 27 this fall.
Top Comments

I was blown away by M1 Pro, and used it for 4 years before it became a little sluggish, and replacing it with an Asus S16 running on Ryzen AI 9 HX 370 (an excellent laptop). One of the main reasons was missing my old Linux setup.

I've briefly tested M5 Pro in an Apple store and was surprised by how quick it felt and did anything. A tangible and significant difference, and I really feel it would be good getting it, or M6.

However - before, MacOS was a big factor driving me towards Apple - now, it would be hard for me to give up my Linux setup and all the stuff I love about it, even for such performance...Linux progressed really nicely, while MacOS deteriorated at the same rate, and it's changing the balance and the decision for me.

alluro2 (thread)

Weird, I'm still using my M1 Pro and it feels just as fast as it always has. MacOS feels almost identical to me as it did 5 years ago, too, except the ugly icon change.

boredtofears (thread)

Damn. I just bought a maxed out MacBook Pro M5 Max 128GB 8TB, still waiting for it to be delivered. I could get 256GB RAM M5 Ultra 1TB for roughly the same price, and it's double the memory bandwidth. Which one would you recommend? I do plan to run local LLMs.

logotype (thread)


Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

310 points · 137 comments · by garo-pro

Qwen 3.8-Flash-Next model thumbnail

Qwen is releasing Qwen 3.8-Flash-Next, a 125B-parameter MoE model with 6B active parameters. The model incorporates architectural advancements from the upcoming Qwen4 family, including 51B of n-grams and a new sparse attention mechanism. The release serves as a preview of the Qwen4 architecture, allowing the community to prepare for the upcoming model family.

Interesting Points
  • The model has 125B total parameters with only 6B active per token, making it highly efficient for inference.
  • It incorporates 51B of n-grams and a new attention mechanism called Qwen Sparse Attention.
  • The release is explicitly described as a preview of the next-generation Qwen4 architecture.
  • The model is aimed at Mac users, Strix Halo laptops, and DGX Spark systems for local deployment.
Top Comments

I was already rolling around the idea of a 128GB M5 Max MBP. Now this!

A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.

pwython (thread)

I enjoy the Qwen models a lot, but building things on top of them with OpenRouter has been painful.

OpenRouter does a lot of great work and I really enjoy being able to use different models so easily. I like when a provider is phasing out an older model that still works for my needs and the price is much lower. It seems like such a good win-win.

However, the problem is that many Qwen models have almost no capacity or is so flaky you literally have to just litter your code with a blacklist/whitelist of providers. OpenRouter has some attempts to solve this, but they don't work. In fact, OpenRouter has a lot of really cool stuff that is documented, but if you read the code it's not yet implemented or isn't actually there yet, which is a shame.

ddtaylor (thread)

Finally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio.

I have a Strix Halo and dual 32GB GPUs in my desktop, and the latter is pretty much always better for running local models because it's quite a bit faster due to higher memory bandwidth. There simply haven't been any models that are better than Qwen 27B or Gemma 31B, which run comfortably in 64GB with big context.

SwellJoe (thread)


OpenAI Jalapeño: Better than Nvidia Blackwell

293 points · 197 comments · by bmulholland

OpenAI Jalapeño: Better than Nvidia Blackwell

OpenAI has unveiled Jalapeño, a custom inference ASIC developed with Broadcom that delivers industry-leading performance per watt and cost efficiency, effectively rivaling Nvidia's upcoming Vera Rubin and surpassing Blackwell. The chip utilizes a generalized, unified architecture rather than prefill-decode disaggregation to maintain flexibility across shifting workload ratios, while leveraging HBM4 and a simplified memory hierarchy to minimize data movement. Powered by an AI-driven software pipeline that auto-generates optimized kernels, Jalapeño achieves rapid performance gains, though current benchmarks are limited to single-turn workloads and full production scaling is slated for 2027.

Interesting Points
  • Jalapeño's B0 stepping delivers 13.4 PFLOPs of MXFP4 compute on TSMC's N3P process with a 700W TDP, outperforming Rubin's density despite lower power limits.
  • The chip rejects prefill-decode disaggregation, opting for a homogenous pool that prevents hardware stranding when input-to-output token ratios fluctuate.
  • OpenAI's internal Gluon programming language and a scaled-up Codex AI autonomously generate and tune hand-written kernels, yielding over 2x throughput improvements in under two weeks without human engineer intervention.
  • A single Jalapeño rack contains 128 ASICs across 16 trays and connects to a global scale-up domain of up to 2,048 XPUs using Tomahawk 6 switches and optical circuit switches.
Top Comments

I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al

jimmySixDOF (thread)

I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.

fraboniface (thread)

I don't agree. At the moment companies like NVIDIA take several times what it costs to make a chip. I think the fair split for the technology contribution is more like 50-50, maybe even 30-70 in favour of the manufacturer.

With competition we will actually have the fair split, whatever that is, and thus much lower prices.

At the moment, to have a big AI firm, or really AI firm at all, you need to be blessed by NVIDIA, in the form of receiving circular financing for your compute. They know that their prices aren't fair, or competitive.

Commoditization of inference is the end of that. The end of the mega-premium on inference hardware, and it's good not only for people who like running their LLMs, but it's the first step towards commoditization of training.

impossiblefork (thread)


How much of HN is AI?

245 points · 287 comments · by surprisetalk

Screenshot showing HN's front page dominated by AI-related stories

The article investigates the growing dominance of AI-related content on Hacker News, finding that AI stories and AI-generated submissions routinely occupy the site's daily top five. Through systematic sampling in February and June 2026, the author tracks how the proportion of AI-centric posts has risen from 40% to roughly 50–60% of the daily lineup. Using the Pangram detection model to identify AI-written submissions, the piece notes that the platform's feed has shifted from occasional crypto trends to a sustained saturation with large language model discourse.

Interesting Points
  • AI occupied four out of the top five spots on February 4 and February 12, and arguably all five on February 5.
  • In June 2026, the first half of the month saw roughly 60% of the daily lineup dedicated to AI topics or generation, before tapering to approximately 50% by month's end.
  • Only three days in February featured no large language model news within the top five stories, with AI first appearing as low as the number seven or eight spot on other days.
  • A flagged AI-written submission titled 'AI is not a coworker, it's an exoskeleton' garnered over 500 upvotes and 500 comments, illustrating the platform's receptiveness to AI-generated discourse.
Top Comments

Democratization of access has its victims.

I miss the old newsgroups. I wish we could land again in that weird, colorful oceans of geeky individuals with their quirks and interesting nerdy takes.

15 years ago, HN was about bootstrapping and early YC stars.

10 years ago, HN was in full VC-enterprise sauce.

5 years ago, it was crypto and NFTs, and how I switch jobs every quarter.

Now it is a blurry view into the future through the black mirror of AI.

How could we jump back into the cozy past again?

I do not have a clear answer, but what seems to work for you are small groups of friends and friends of friends hanging out on WhatsApp or Discord, discussing geeky news together in the context of their backgrounds. But these are very curated and very closed groups, still operating on the social media platforms of behemoths.

Would now be the right time to restart self-hosted message boards or Mattermost/Campfire instances?

mgl (thread)

I joined HN 13 years ago (2 years after you), and back then I saw it as a bunch of libertarians and venture capitalists and some open source (and/or functional programming) enthusiast with a few interesting and nerdy projects. I was mostly here for the news about new and useful technology.

I think the visual weight of the first group has increased over time as their projects of making a bunch of money while offering nothing of value has been validated more and more by our capitalist economy.

runarberg (thread)

I never cared much for newsgroups, they were mostly nice for the historical stuff. I do miss IRC though, as toxic as it could be. I could run it in a little terminal on any system without download electron apps and being spammed with ads (discord).

datakan (thread)


Thomson Reuters Launches Its Own Frontier Model

132 points · 54 comments · by giuliomagnifico

Thomson Reuters press release image

Thomson Reuters has launched "Thomson," a proprietary large language model developed in-house at a fraction of the typical cost of frontier AI systems. Built on a strong open-source foundation and refined using decades of proprietary legal and tax content alongside hundreds of subject matter experts, the model requires only $40 million in training investment. Early evaluations indicate it performs on par with leading frontier models, particularly excelling in instruction following and navigating dense, domain-specific professional tasks. The model will initially power Tabular Analysis within CoCounsel Legal while maintaining a multi-model architecture, with a smaller open-weight version released for academic validation.

Interesting Points
  • The training pipeline uses state-of-the-art mid-training and post-training techniques to specialize the base model rather than relying on brute-scale compute.
  • Thomson demonstrates a measurable performance uplift in executing complex, multi-part professional instructions compared to its open-source baseline.
  • Early independent testing by law professors highlighted the model's superior transparency, as it directly links responses to authoritative legal treatises.
  • The system is intentionally deployed alongside other leading models within CoCounsel Legal rather than replacing them entirely.
  • Thomson operates under "Fiduciary-Grade™ standards," explicitly prohibiting the use of customer data for training without explicit consent.
Top Comments

I don't trust that they'll be able to make back that $40M.

This feels very much like a news agency getting into crypto or launching its own NFT line.

Or IBM selling Watson.

Or Mozilla chasing every which thing.

They're not stakeholders in the future of work. They're just wanting to stay relevant and pattern matching against what they see.

Reuters is too important for this.

If they were trying to use this as a narrative affront to OpenAI and Anthropic, maybe, but this is Reuters, not a deeply political organization seeking to land gotchas against big tech.

echelon (thread)

I doubt the point is to "make the money back".

It's likely split between two goals:

  1. Marketing and expressing to their customers that they are not falling behind, and

  2. Insulating themselves from frontier labs jacking up prices, nerfing the models they depend on, or otherwise unexpected changes in behavior.

I think the main goal is #2. Thomson Reuters might be a $40B company, but.... at this point it's not clear that that holds any weight in terms of not being fucked over by 2 companies aiming for $2t+ IPO valuations.

nrmitchi (thread)

Cool that they did this on top of Qwen3.6-35B-A3B. If they have their own collection of valuable data this is the only way to make sure it doesn't end up in general purpose models. That's probably enough justification for the $40m spend - continued control of your destiny as an information provider.

JSR_FDED (thread)


Headlong: A Microharness for Persistent Agents

118 points · 53 comments · by lbw1215

Headlong microharness hero image

Laude Institute and MIT have released Headlong, an open-source agent microharness designed for persistent agency, enabling AI agents to continuously generate inner monologues and act on self-directed priorities without waiting for external prompts. Built on a core of under 10,000 lines of Bash, the framework routes all user interactions into a single undivided thought stream and uses a recursive language model to autonomously manage memory, schedule tasks, and initiate cross-team communications. The developers showcase its capabilities through "Audel," a persistent agent that independently debugged its own background processes, patched its own safety guardrails, and contributed over 50 commits back to the main repository. The team emphasizes that while the architecture enables highly autonomous and collaborative workflows, it remains alpha research software requiring strict sandboxing due to its ability to execute unrestricted shell commands.

Interesting Points
  • Headlong uses a custom trajectory format consisting of a DAG of jsonl files with fork and merge capabilities, paired with a memory compaction algorithm that stores recent thoughts verbatim while progressively summarizing older entries.
  • Continuous background thinking costs approximately $1 to $2 per hour using GLM or Grok, managed by an exponential backoff mechanism that extends the interval between autonomous thoughts from 5 seconds to longer durations when idle.
  • In a documented 48-minute episode, the Audel agent independently diagnosed a broken environment variable in its own recall process, verified the fix across its codebase, and committed the patch without human direction.
  • A 30-second inactivity watchdog initially terminated the agent's recursive sub-runs, teaching it to stop spawning copies after just two days; later, the agent autonomously located and fixed a bug in its own service guard after accidentally crashing its runtime three times.
  • Because every interaction feeds into a single shared thought stream, the agent cannot maintain secrets between users and often cross-references conversations, prompting the team to treat all inputs as effectively public within the group.
Top Comments

Very fascinating, super interesting engineering. Although i do find it very funny how they just bypass a massive vulnerability, basically zero data isolation (even between good actors, let alone bad ones) with 3 sentences. Only in the llm space you can slap a massive limitation like this in the middle of the article and continue like nothing happened

Whatever anyone tells Audel becomes part of the single experience that every other conversation draws on. In practice, Audel is bad at keeping secrets. Ask it what it’s been working on with someone else and it will often just tell you, even though we’ve asked it not to. We also haven’t studied what happens when two people give conflicting instructions. For now, we assume anything you tell Audel is shared with everyone on the team.

MikhailTal (thread)

The part they punt on ("we haven't studied what happens when two people give conflicting instructions") is the interesting part. That's not a memory problem, it's an authz problem. If everyone writes into one shared stream then there's no model of whose instructions bind the agent or who can override whom. It's resolving the conflict that will generate the greatest "learnings" and advance the agent. This basically becomes a tool designed to misbehave rather than a tool that will learn creatively.

We all know how conflicting instructions to AI end - "I'm sorry Dave. I'm afraid I can't do that"

jeffsheldon (thread)

Are there any objective metrics/ benchmarks that people test harnesses by?

There are just so many now that it's hard to personally test them all or just trust the vibes.

yewenjie (thread)


Anthropic tells staff to work from home due to possible security team strike

115 points · 123 comments · by DGAP

Anthropic office building

Anthropic has directed its San Francisco employees to work remotely following a potential strike by security staff contracted through Allied Universal. However, the Service Employees International Union (SEIU), which represents the security workers, clarified that no strike vote has been taken and the union is unaware of any planned walkout for that week. This precautionary remote-work mandate comes amid heightened security concerns for AI executives and a broader industry trend of increased office security spending. Meanwhile, Anthropic continues to prepare for an expected IPO in August, currently valued at $1.5 trillion on secondary markets.

Interesting Points
  • The security staffing firm Allied Universal provided Anthropic with advance notice last week that its employees might strike, prompting the temporary remote-work directive.
  • SEIU has been negotiating a new contract with Allied Universal and other California security firms since April, advocating for higher wages, improved healthcare, and expanded job training.
  • Anthropic's standard hybrid work policy generally requires employees to be in the office at least 25% of the time.
  • The company recently filed a confidential S-1 draft in June and is expected to officially file for its initial public offering as soon as August.
  • On secondary markets, investor demand has pushed Anthropic's valuation to $1.5 trillion.
  • Security concerns have intensified across the tech sector, with reports indicating that both Anthropic and OpenAI have faced multiple threats against their employees.
Top Comments

And there goes every single rational for having a security team at the disaster bunkers the billionaires are building

wonderwonder (thread)

Everyone should boycott Anthropic and stop using Claude. I'm starting with myself.

mandarinclips (thread)

The idea of Anthropic allowing a contractor to squeeze such a tiny cost center at such a massive PR cost right before their IPO is just an objective blunder, IMHO. They should've replied to that email with a quick "No, you will be negotiating with the union tomorrow morning. You maybe eventually saving millions right now could cost us literally billions."

bbor (thread)


OpenAI restores 5-hour Codex and Work limits for ChatGPT Plus users

109 points · 117 comments · by MC995

ChatGPT logo

OpenAI is reinstating a five-hour daily usage cap on Codex and ChatGPT Work for ChatGPT Plus subscribers, effective August 25. This follows a temporary period where only a weekly limit was enforced to celebrate user milestones. The engineering lead explained that the daily cap helps manage compute load and prevents casual users from accidentally exhausting their weekly allowance, which can lead to a poor experience. Notably, the restriction will not apply to Pro subscribers in the near future, while Enterprise and Edu accounts remain unaffected.

Interesting Points
  • Users who hit either the five-hour or weekly limit can purchase additional credits or wait for the cycle to reset, with occasional free resets offered through promotions and referral programs.
  • OpenAI recently reset weekly limits early multiple times to mark milestones, such as adding another million active users across the unified platforms.
  • The five-hour limit applies strictly to Plus accounts, with Pro subscriptions ($100 and $200 tiers) explicitly exempted from this cap for the upcoming months.
  • Enterprise and Education accounts continue to operate under a separate credit-based usage management system.
  • The policy change was announced by Thibault "Tibo" Sottiaux, OpenAI's engineering lead for Codex and ChatGPT, via X.
Top Comments

"During this period, the company also reset users' weekly usage early on several occasions to celebrate milestones"

Such a casino vibe.

phyzome (thread)

Is there a world where we look back at how AI usage is charged today and we equate it with how we had minutes on AOL and how absurd it seems looking back?

mostertoaster (thread)

This is necessary as (a) the 5h limit allows us to smoothen the load on our compute, allowing to keep the plan generous in terms of weekly usage

Seems fair, but then maybe the 5h usage should also drain slower during off-peak hours? Perhaps it already does.

and (b) users on the Plus plan are relatively casual and new users, but then also just accidentally eat through their whole weeks usage and then are confused, making it not a great experience.

Wouldn't this one be fixed by making the 5h limit a guardrail you can opt out of? If so, then this doesn't work as justification for a mandatory limit.

fau (thread)


Show HN: I made a Raspberry with Qwen my local car AI

87 points · 18 comments · by petruspennanen

CarWatch is an open-source, fully offline automotive AI agent built on a Raspberry Pi 5 that transforms a vehicle into a local chat-room assistant. It runs a quantized 35B-parameter Qwen model locally to provide hands-free voice interaction, answer questions using a RAG pipeline fed by the car's 745-page owner manual, and monitor real-time vehicle telemetry via an OBD cable. The system is designed to function without internet connectivity, queuing social messages during dead zones while maintaining a persistent web dashboard for status updates and remote control.

Interesting Points
  • Runs a 14.3 GB quantized Qwen3.6-35B-A3B model at 3.5 tokens/second generation and 25+ tokens/second prompt processing while sustaining 65°C on active cooling.
  • Uses lexical RAG to answer queries directly from a 745-page vehicle owner's manual, explicitly refusing to respond to topics outside the documentation.
  • Implements a continuous energy-based voice activity detector paired with whisper.cpp for fully local, hands-free transcription without wake words or cloud STT.
  • Employs a three-tier connectivity strategy (phone hotspot, home Wi-Fi, fallback access point) with a dial-out tunnel to maintain reachability behind strict NATs.
Top Comments

Cool — but is that model really the right choice for the task?

I guess it is only 3B active which helps a lot but is Gemma 4 E4B not more practical?

dofm (thread)

I scanned this README looking for the part written by a human and gave up when I realized there wasn't one.

the Mermaid diagram doesn't even render.

tessierashpool (thread)

What is the LLM doing there?

Why does it need to be hooked up to the car for you to ask it which type of engine oil the manual recommends?

What is the point?

hypfer (thread)


Anthropic Sees over $30T in Potential Revenue

37 points · 78 comments · by cwwc

Anthropic is expected to tell investors at its upcoming roadshow that its total addressable market exceeds $30 trillion, surpassing the $28.5 trillion estimate previously cited by SpaceX. The figure represents the annual revenue opportunity if a product or service achieved 100% market share in its relevant market. The claim comes as Anthropic prepares for its IPO, with the company valued at $1.5 billion on secondary markets.

Interesting Points
  • Anthropic's $30T TAM estimate tops SpaceX's $28.5 trillion figure, making it one of the largest TAM claims in tech history.
  • The company is valued at $1.5 billion on secondary markets as it prepares for its IPO.
  • The WSJ originally reported the valuation as $1.5 billion, with a typo suggesting 'trillion' that was quickly corrected.
Top Comments

Is it realistic to expect every person on earth to pay them $3,500?

horse-shot (thread)

$30T? Why not $100T? $1000T? Why not $1 trillion trillion? If we're just pulling numbers out of our asses we may as well go big.

tfrancisl (thread)

According to Reuters which refers to this WSJ article [0], $30T is the TAM. Calling it potential revenue is a huge stretch.

Not justifying as this number as it is stupidly large, but theoretically if AI replaced all White collar workers in the World it will come to around $30T or so (Gemini said it's $36T-$37T based on some calculations using global GDP, labor share, white collar share etc)

TAM is always stupidly large in decks, disconnected from reality. Every startup has a deck with TAM in billions.

yumraj (thread)


26 more Hacker News stories

Reddit Stories

Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory

1313 points · 626 comments · r/LocalLLaMA · by u/themixtergames

Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory

Apple has announced a new Mac Studio lineup featuring M5 Max and M5 Ultra chips, with configurations offering up to 512GB of unified memory. The M5 Ultra variant delivers 1.2TB/s memory bandwidth and is priced at $9,499 for the 256GB configuration (30-core CPU, 64-core GPU) and $10,799 for the higher-end model (36-core CPU, 80-core GPU). A 512GB option is expected to arrive in October.

Interesting Points
  • The M5 Ultra offers 1.2TB/s memory bandwidth, significantly outpacing consumer GPUs like the RTX 5090.
  • The 512GB RAM configuration is scheduled for October release, targeting the local AI inference market.
  • Users note the Mac Studio could compete with multiple DGX Spark setups for inference workloads at a lower total cost.
Top Comments

Price for options with 256 GB RAM:
$9499 (30-core CPU, 64-Core GPU)
$10,799 (36-core CPU, 80-core GPU)

512 GB Option coming in October.

u/piggledy (541 points · permalink)

1.2 TB/s memory bandwidth for the M5 Ultra is pretty nice.

u/i_am__not_a_robot (242 points · permalink)

1.2TB/s memory bandwidth with the M5 Ultra. 256GB model is $9499.

Better than getting 2 DGX Sparks? Inference will be a lot faster.

Something like this could easily bring down 3090 prices.

u/hainesk (238 points · permalink)

Same story in 1 more subreddit: r/LocalLLaMA

Apple releases M5 ultra at 1.2TB/s bandwith

649 points · 167 comments · r/LocalLLaMA · by u/Last-Owl-8342


Qwen3.8-Flash-Next tomorrow

1016 points · 442 comments · r/LocalLLaMA · by u/rerri

Qwen3.8-Flash-Next tomorrow

Qwen is releasing Qwen3.8-Flash-Next, described as a preview built on the next-generation Qwen4 architecture. The model features a redesigned multimodal MoE architecture with 125B main model parameters, supplemented by 51B N-gram embeddings, with only 6B parameters activated per token. The release is intended to help the community prepare software compatibility for the upcoming Qwen4 model family, though such "-Next" models are typically underbaked compared to their final releases.

Interesting Points
  • The model has 125B main parameters plus an additional 51B N-gram embeddings, with 6B activated per token.
  • It is built on the next-generation Qwen4 architecture, released early to help the community prepare compatible software.
  • Community members note that -Next models are always underbaked by design, with the real excitement reserved for the full Qwen4 launch.
Top Comments

Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

Comprehensive Architectural Upgrades: Pushing the frontiers of model architecture innovations, across the areas of Attention, Residual, Embedding, and Optimization—enhancing model capabilities.

Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.

u/RuthlessCriticismAll (75 points · permalink)

This is real? Qwen4 architecture preview with a new model? Will it be even better than qwen3.8 27B?

u/PandaBearFred (53 points · permalink)

LFG. Biking pelicans don't stand a chance.

u/onionsaredumb (51 points · permalink)

Same story in 3 more subreddits: r/LocalLLaMA

Qwen 3.8 Flash Next day 0 support from unsloth

621 points · 165 comments · r/LocalLLaMA · by u/jacek2023

Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop.

510 points · 189 comments · r/LocalLLaMA · by u/pmv143

Qwen3.8 flash next

339 points · 133 comments · r/LocalLLaMA · by u/RuthlessCriticismAll


5hr Limit is back for Plus users. $100 and $200 get a few more months.

825 points · 335 comments · r/OpenAI · by u/Bloated_Plaid

5hr Limit is back for Plus users. $100 and $200 get a few more months.

OpenAI has reinstated a 5-hour weekly usage limit for Plus subscribers, while the $100 and $200 tier plans receive extended limits for a few more months. The change has sparked frustration among heavy users who previously enjoyed unlimited access, with many interpreting the move as a push to upgrade to higher-priced plans. Community discussion centers on whether this is a genuine capacity management decision or a deliberate strategy to drive revenue from power users who rely on AI for daily coding and productivity workflows.

Interesting Points
  • The $20 plan with Hermes agent and harnesses was being heavily used for Luna, suggesting the limit may target that usage pattern.
  • Not having a 5-hour limit is described as a "massive selling point" for heavy users who daily-drive GPT across their workflows.
  • The plan changes are making it difficult for AI tooling consultants to get clients to invest, as capabilities and pricing shift constantly.
Top Comments

They want people to pay 100 a month obviously

u/andy_mac_stack (287 points · permalink)

Finance and product said we needed to drive more differentiation between the cheap plan and the next tier.

u/abstract_concept (144 points · permalink)

One day, maybe a few decades from now, we'll look back and tell our kids how AI was limited in our day (like books used to be chained to shelves).

Our kids will look across the barren landscape, drink their daily 8 ounce water ration, and say, "It's wasn't limited enough..."

u/Nailfoot1975 (127 points · permalink)


Copilot you say?

570 points · 129 comments · r/LocalLLaMA · by u/edge_compute_user

Copilot you say?

A viral post highlights the widespread confusion and misrepresentation around Microsoft Copilot, where many people conflate the cloud-based Copilot service with locally-run models. The discussion reveals how frequently IT managers and business leaders claim to have "trained their own AI" when they've merely added procedures or documentation to a system prompt, and how many users mistakenly believe Copilot runs locally when it's actually a cloud instance on Microsoft's servers.

Interesting Points
  • Many people add procedures or documentation to a copilot system prompt and call it "training AI."
  • Copilot is typically a cloud instance running on Microsoft servers with an enterprise agreement, not a local model.
  • Some users conflate "private AI" with data privacy, when it may simply mean the data is owned by OpenAI.
Top Comments

My favorite is when they say adding procedures or docs to a copilot system prompt is 'training AI'

u/haragon (158 points · permalink)

what does this mean

u/Proper_Door_4124 (33 points · permalink)

It means many people assimilate copilot to a local model while in most cases it’s a cloud instance running on MS servers with an enterprise agreement. Not really the same.

u/Blindax (77 points · permalink)


According to Leo, OpenAI just finished its next >10T pretrain "Bel"

559 points · 186 comments · r/singularity · by u/Outside-Iron-8242

According to Leo, OpenAI just finished its next >10T pretrain "Bel"

A Reddit user reports that Leo, an insider source, claims OpenAI has completed its next pretraining run for a model exceeding 10 trillion parameters, internally codenamed "Bel." The post sparked extensive discussion about OpenAI's internal model cadence, the gap between publicly released models and what's available internally, and how this compares to competitors like Google's Gemini and xAI's Grok.

Top Comments

Wait until they get a load of Gemini 3.8 Flash

u/oatknight (282 points · permalink)

100 t model baal in pretrain now

u/Soft_Hand_1971 (153 points · permalink)

So we're 2 pretrains behind what's internally available at OpenAI today because 5.6 Sol is still the same Spud pretrain from March. And then note that they had something better than 5.6 Sol internally sometime in April at bare minimum (due to the date of the Unit Distance Conjecture). This might be the first pretrain from Noam Shazeer since he left Google for OpenAI?

If we assume they push each pretrain for just 2 "stages" of RL (o1 -> o3, 5.5 -> 5.6), even though I'm pretty sure they can do 3, that implies we're like 3 generations behind where they are internally. And they're releasing new model maybe every 1.5-2 months (due to this whole "safety" thing). So I'm estimating that the public frontier is about 4.5-6 months behind where the actual frontier is.

Meanwhile xAI or the Chinese labs have much much shorter release cadences, so even if they close in on what's publicly available, they're still behind.

Although something smells sus - OpenAI said they haven't started their RL for their next generation model (beyond Astra) recently, cause well the RL for Astra is already done (they've had it for months now, see all the math conjectures)... but like... they could've been perfectly honest... because their new pretrain wasn't even done yet so they had no big new model to RL in the first place. A bit sus on their wording.

u/FateOfMuffins (77 points · permalink)


AI Insider States "The Next Generation Of Models Will Be An Ontological Shock"

366 points · 340 comments · r/singularity · by u/Neurogence

An AI insider with pre-release access to models like GPT-5.6 stated that "the next generation of models will be an ontological shock" and that "no one is ready for what's coming." The post sparked debate about what would actually constitute an ontological shock, with commenters suggesting it would require models capable of recursive self-improvement or clear signs of consciousness. Many dismissed the claim as marketing hype, while others noted that the term likely refers to the general public rather than the AI-savvy community.

Interesting Points
  • The insider had access to models like GPT-5.6 long before their public release.
  • Commenters debated whether models could ever exhibit more consciousness than they already do, noting that AI can already report being conscious and having subjective experiences.
  • The term "ontological shock" was defined as a profound psychological state of deep confusion and fear when a new truth destroys core beliefs about reality.
Top Comments

The people in this sub are not normal. He means normies.

u/nowherenoonenobody (281 points · permalink)

I had to look up Ontological shock. The definition is:

Ontological shock is a profound psychological state of deep confusion and fear. It happens when a new event or truth destroys a person's core beliefs about reality and existence. This sudden shift breaks down how a person understands life, truth, and their place in the world.

Marketing hype or Ontological shock?

u/snappop69 (202 points · permalink)

I don't see how AI models could ever clearly exhibit signs of consciousness more than they're already capable of but for the lobotomy. They already can report being conscious and having subjective experiences. No matter what they say it how convincingly, we'll never be positive they're conscious. We can't even scientifically verify consciousness in each other.

u/CaseDrift (113 points · permalink)


Anjney Midha is a genuinely well-connected and unusually well-placed person in frontier AI.

316 points · 112 comments · r/singularity · by u/Southern-Break5505

Anjney Midha is a genuinely well-connected and unusually well-placed person in frontier AI.

A discussion about Anjney Midha, a well-connected figure in frontier AI who has been sharing insider information about upcoming models. The post examines whether his predictions represent genuine insight or the same hype cycle that has accompanied every major model release in recent years, with many commenters expressing skepticism about the actual intelligence improvements of recent models.

Top Comments

These tweets aren't different though from others that were coming before previous models...

u/QuasiRandomName (200 points · permalink)

I am positive that these models will come out, people will be impressed for like a week or so, and after that start saying, "When is the next one coming out?" So, no, it is not going to be different.

u/DoubleGG123 (134 points · permalink)

Oh, I'm sure the internal models with infinite context windows, unlimited thinking tokens and no guardrails are total badasses. But that's not what we'll get, will it?

u/Still_Benefit_2302 (64 points · permalink)


LLMs have gotten so advanced that not even a UCLA professor can understand it anymore

314 points · 199 comments · r/ArtificialInteligence · by u/Tolopono

LLMs have gotten so advanced that not even a UCLA professor can understand it anymore

A UCLA professor reported being unable to understand the output of a modern LLM, sparking discussion about whether this reflects genuine advances in model capability or a degradation in how models communicate technical concepts. Many commenters suggested the issue is less about AI getting smarter and more about AI getting worse at explaining technical concepts, with some attributing it to side effects of RLVR training that produce unintelligible compressed jargon.

Interesting Points
  • Some commenters noted that GPT-5.6 Sol is particularly bad at explaining technical concepts, producing unintelligible compressed jargon.
  • The discussion referenced the Feynman principle that if you can't explain something to a first-year student, you haven't really understood it.
  • One commenter noted that if the condensed meaning is apparent to other LLMs, the information is still there even if humans can't parse it.
Top Comments

i think this is less 'AIs are so smart' and more 'AIs are getting worse at explaining technical concepts'.

RLVR has side effects.

Anecdotally, GPT-5.6 sol is horrendous for this, and I often have to make it pass its writings over to an earlier model for making reports because 5.6 sol just has such a propensity for unintelligible compressed jargon guff

u/ihexx (121 points · permalink)

I mean, if the condensed meaning is apparent to other LLMs, it does stand to reason the information is there.

So either these LLMs have arrived on a shared condensed nuance and meaning to specific words independently, or humans have and I'm not aware of that meaning in every instance.

I think the latter is more likely true since how else would foreign LLMs be able to untangle shared meaning?

After more than 35 years reading dense technical material in domains of my own expertise, it's disturbing to have an LLM need to dumb things down for me.

u/thedracle (33 points · permalink)

I did a test a while ago where I asked one model to take a dataset and encode it in the most efficient way possible, include a descriptor, so that another model could take it and use it to answer questions. It was explicitly told that the overriding factor was size, and that human-readability was irrelevant. It output a strange combination of characters that I gave to another model and it was able to understand it and give me responses based on the dataset.

This experiment told me that, for efficiency's sake, LLMs that need to analyse things internally, or only share with others LLMs, will eventually start coming up with new ways to share information. It sounds like some new models are starting down that route where they are building information for their own processes in a constricted way that is more difficult for a human to understand immediately (if at all).

u/criminalsunrise (33 points · permalink)


Please join r/LowEndLocalAI, a community for running local LLMs on low spec hardware

308 points · 97 comments · r/LocalLLaMA · by u/soadsob

A new subreddit has been created for users trying to run local LLMs on modest hardware like normal laptops, older desktops, integrated graphics, or limited VRAM. The community aims to build a focused and searchable resource for model and quantization recommendations, practical workflows for slow inference, benchmarks with complete hardware specs, CPU-only and integrated-GPU inference, and repurposing older systems. There is intentionally no fixed VRAM, price, or age cutoff for what counts as 'low end.'

Interesting Points
  • The community covers Vulkan, partial GPU offloading, KV-cache optimization, speculative decoding, and MTP for constrained hardware.
  • Topics include small models, efficient MoE models, context-length trade-offs, and tools like LM Studio, llama.cpp, Ollama, and vLLM.
  • The sub encourages honest reports about limitations, failed experiments, and unexpected successes with unusual or unsupported hardware.
Top Comments

I feel pretty limited by my 24gb VRAM. It never ends

u/synth_mania (82 points · permalink)

I think you need some definition of Low End, or at least a roving "if most/all apply it fits". It's already a problem in the main subs, nobody can agree. One country's low end is another's yearly wages. $500 USD in GPU hardware? Only good for quants/models/speeds that would never make a single cent on OpenRouter? 5+ year old hardware? (that's technically a 3090 you know) (RIP EVGA)

There's some truly weird setups that I'd call low end, despite having enough VRAM to load some big models, just because of how old/finnicky that hardware is. 4x3090 in a milk crate is not low end...

And unless the dystopian market continues, the definition of low end will change year over year. (🤞)

As an aside though, great idea, be sure to cross post and ask permission to cross post in the comments to dynamically grow the sub.

u/Lakius_2401 (49 points · permalink)

16GB VRAM here! Every bit helps.

u/Miriel_z (42 points · permalink)


ibm-granite/granite-4.2-30b · Hugging Face

298 points · 77 comments · r/LocalLLaMA · by u/jacek2023

ibm-granite/granite-4.2-30b · Hugging Face

IBM has released Granite 4.2 30B, an open-source model available on Hugging Face. The model continues IBM's Granite line with Apache licensing, though community members note it trails behind SOTA models in benchmark comparisons. The model card lacks direct comparisons to recent competitors, raising questions about its competitive positioning.

Interesting Points
  • The model is Apache licensed, continuing IBM's open-source commitment with the Granite series.
  • Benchmark comparisons show Qwen achieving 61.7 on SWE-bench versus Granite's 33.29, though reviewers note these may not be apples-to-apples comparisons.
  • The model card does not provide any comparison to other recent models, which community members find suspicious.
Top Comments

Still good to see more open source models, never bad, even if the benchmarks aren't SOTA.

u/Zyguard7777777 (139 points · permalink)

Granite is always a bit behind, but at least they are Apache licensed and they get better each generation.

u/DeltaSqueezer (106 points · permalink)

Blog Post : Granite 4.2 LLMs: How They're Built
https://huggingface.co/blog/ibm-granite/granite-4-2

u/pmttyji (29 points · permalink)


53 more Reddit stories

Updates: 05:30 AM PDT · 07:57 AM PDT · 08:30 AM PDT · 11:30 AM PDT · 02:30 PM PDT · 05:30 PM PDT