Kimi K3 Reshapes Open AI Race Amid Geopolitical Shifts
Overview
The conversation is overwhelmingly driven by Moonshot’s Kimi K3, whose benchmark-breaking performance and upcoming open-weight release have reignited fierce debates over open-source versus frontier models. Corporate and regulatory tensions are simultaneously escalating, with Apple widening its lawsuit against OpenAI, the EU forcing Google’s data sharing, and the White House rolling out a new frontier AI vetting program. Security risks also grabbed headlines following a sophisticated AI-driven breach at Hugging Face and the deployment of agentic tools for proactive code remediation. Amid this backdrop, Chinese policymakers champion open-source cooperation while US developers and enterprise leaders navigate the rapid practical integration of AI across healthcare, software, and daily workflows.
Hacker News Stories
Blatant AI slop just won a 25k USD DeepMind Kaggle Grand Prize
436 points · 273 comments · by twerkmeister
A submission widely described as AI-generated slop has won the DeepMind Kaggle Grand Prize for measuring AGI, with a $25,000 prize. The controversy centers on the use of AI-generated submissions and AI judges in the competition, with commenters noting the irony of AI submissions being evaluated by AI. The winning paper, titled "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition," has been criticized for containing obvious Claude-generated text, and the broader discussion questions whether hackathons and competitions can survive the influx of AI-generated content when judging is also automated.
Interesting Points
- The winning submission's paper title 'Scale Buys Evaluation but Not Control in AI Metacognition' was identified by commenters as containing a distinctive Claude-generated phrase pattern.
- Commenters noted the competition used AI judges to assess submissions, creating a loop where AI submissions are evaluated by AI evaluators.
- One commenter observed that the winning submission had a strong 'stench of AI about it, across multiple participants' in the discussion thread.
- The broader concern raised is that hackathons and competitions are being 'killed by AI' because all code is generated and judging happens via AI, with insiders as the main winners.
Top Comments
onesandofgrain (7 replies)
AI is 95% useless. Not quite worth the trillion dollar market cap lol.
- The AI bots are downvoting me * hooray
hoppp (7 replies)
I don't know about this exact competition but overall fair hackathons have been killed by AI.
It all seems fine from the outside but all the code is generated in all the projects and judging happens via AI, I have seen projects win because they prompt inject that they are the winners.
It used to be about human skill, now it's about ideas and of course insiders are the main winners.
ecshafer (5 replies)
AI is useful. But the amount of people that are simply offloading all of their thinking to AI and blindly accepting the answer is absurd. Kaggle is most likely using ai to assess the submissions and are not using any common sense by blindly accepting the results.
throwfaraway135 (2 replies)
AI submissions and AI judges a match made in (AI) heaven.
ablation (4 replies)
"I think you just need to accept the results of the competition. The winning submissions clearly provide value and had a lot of effort invested in them. I'm not really worried about a few inconsistencies or mistakes if the value is still there. Did you think another submission deserved to win over these?"
That comment is gold. Yeah, I'm not worried about hallucinated slop, just accept it was the winner folks.
Apple targets dozens of OpenAI employees with legal letters
370 points · 319 comments · by merksittich
Apple has expanded its trade secret lawsuit against OpenAI by sending legal preservation letters to approximately 40 former employees now working at the AI company. The letters require recipients to retain documents and communications related to allegations of misappropriated hardware engineering and product development secrets. This follows a lawsuit filed last week targeting over 400 former Apple staff now at OpenAI, particularly focusing on key executives like Chief Hardware Officer Tang Tan and senior electrical engineer Chang Liu. OpenAI has publicly denied the allegations, stating it sees no merit in the complaint.
Interesting Points
- Apple's complaint specifically alleges a coordinated effort to access proprietary manufacturing processes and confidential hardware engineering data.
- Tang Tan is identified as a 24-year Apple veteran who previously led product design before transitioning to OpenAI.
- Apple is pursuing individual breach of contract lawsuits against both Tang Tan and Chang Liu for violating their prior employment agreements.
- The company's legal filing characterizes the evidence uncovered during its initial investigation as merely the 'tip of the iceberg.'
- OpenAI's formal denial was issued directly to Bloomberg, where the company explicitly stated it lacks awareness of any evidence supporting the lawsuit.
Top Comments
Zigurd (22 replies)
Everybody wants a platform but nobody wants to spend what it takes to make a platform. That includes things like Windows Phone, Fire Phone, all the glasses, Humane, etc.
As much as everybody hates on OpenAI for chaotic management, they did buy Jony Ive and are presumably giving him everything he wants to build a platform for them. Even though it probably only buys them a 20% chance of success, they haven't doomed the project by underestimating what it takes budget-wise.
And they blew it. Maybe they blew it by not realizing that even long time Apple employees could get arrogant about security. Or maybe it was a loose ethical environment in general. Whatever is it the root or the problem, they set billions of dollars on fire maybe tens of billions, by being unnecessarily cute about Apple proprietary information when they could've been above reproach. They had the resources to hire all the right people with the right knowledge and probably already had them on board.
bix6 (6 replies)
Predictions on who wins? Does Apple actually have a winnable case or are they just throwing a wrench in things?
reenorap (5 replies)
Apple must have hard evidence on this. I can't believe they would take it this far without already knowing they are going to win. If they have to fire a huge chunk of their hardware employees it's going to throw their IPO plans into chaos.
deepwoods (4 replies)
FT frames this as some aggressive escalation tactic, but document retention letters are extremely standard practice. At this point they're basically a formality, as any former Apple employee at OpenAI really ought to know by now that they could get dragged into this. Hold letters can be aggressive if you send them before you've even filed a complaint, but if anything, Apple is late to the party with these.
pembrook (4 replies)
I'm a huge fan of Apple but this kind of thing leaves a bad taste in my mouth.
Regardless of whether OpenAI poached some of their talent or is the one in the wrong, Apple has such a massively dominant hardware business (some might say monopoly level in some areas) that for them to be publicly acknowledging how scared they are of OpenAI…it's just…pathetic.
They're a $5T company and can't muster up the motivation to get in the game and compete in the next computing frontier.
Mozilla: The state of open source AI
358 points · 260 comments · by rellem
Mozilla's July 2026 report argues that open-weight AI has successfully closed the capability gap with closed models and driven inference costs down 50x, making open models the dominant choice for production token routing. However, the competitive frontier has shifted from raw model weights to the agentic harness layer, where closed labs are tightly integrating their scaffolds to create vendor lock-in. The report emphasizes that open AI's survival depends on standardizing governance, solving the portable permission problem for agent actions, and seizing the harness layer before closed ecosystems weld models and orchestration into a single rented product.
Interesting Points
- Inference costs for GPT-4-class models have fallen 50x in 36 months, dropping from $20 to $0.40 per million tokens, outpacing historical computing price curves.
- While 79% of AI developers adopt open models, only 51% reach production versus 63% for closed models, with churn driven by operational hurdles like infrastructure costs, security compliance, and deployment complexity.
- Chinese open-weight models captured over 45% of weekly OpenRouter traffic by April 2026, with Qwen alone accounting for more downloads than the next eight organizations combined.
- A May 2026 benchmark revealed a 21.8-point performance gap favoring a third-party agentic harness over a closed lab's own tooling, but closed providers compressed this to roughly 3 points within eight weeks by tightly integrating their own scaffolds.
- The ecosystem faces an unsolved "write surface" problem, as current protocols like MCP and A2A standardize authentication but lack a portable, cross-framework permission model to safely govern irreversible agent actions.
Top Comments
babblingfish (27 replies)
Speculation: open models is what will kill Anthropic and OpenAI. Hyperscalers can run the models without a licensing fee. Apple can make them smaller and put them on the device.
The frontier models are an edge and a liability. They're astronomically expensive to train. Without them, their models will fade into obscurity. Their marketing depends on people believing the models are meaningfully different, as people have sweatily argued on this forum. Personally, I'm not convinced there's much of a difference between these models at this point. The harness is what takes these random and hallucinogenic models and make them into something deterministic and useful.
GodelNumbering (6 replies)
Exactly 4 months ago, the marketshare on openrouter was 60%-40% in favor of closed models. Now it's 63%-37% in favor of open models. On March 19th, the open models processed 888B tokens in aggregate, yesterday, they processed 4.19T tokens in aggregate. That's almost 5x in 4 months! I can't think of the right intensifier to describe this level of growth.
If you are looking for more details (as inferred by openrouter data), I built a dashboard that updates daily: https://dirac.run/labs-market-share
mft_ (7 replies)
Open models are probably also comparatively astronomically expensive to train - just less so than the frontier models because they're somewhat smaller, +/- the creators are more incentivised to focus on getting more from less compute because they're have to, +/- they rely on distillation of the frontier models and this is more efficient.
But efficiencies aside; creation of open models still requires a lot of money and compute from a large organisation which is willing to accept zero return for that spend. This largesse is unlikely to continue forever; so the question is which will crack first, the frontier models' business model or the fast followers' generosity?
inigyou (3 replies)
There isn't any open-source AI. There is Open AI (not to be confused with the closed company called OpenAI, which was unable to trademark its name). There's no open source AI both because the open source community doesn't have the resources to train a useful AI and because AI doesn't have source code.
brunooliv (3 replies)
This is really insane to me.
There's nothing practical about open-source models yet that makes them even remotely comparable to closed frontier models.
All the hype around GLM, Qwen, now Kimi.... Are people really this naive that they believe these reports or, more worringly, are people NOT using these models and seeing the HUGE gap that still exists?
Kaiser nurses say AI, workplace surveillance are making their jobs, care worse
183 points · 126 comments · by gnabgib
Kaiser Permanente advice nurses report that workplace surveillance and AI-driven performance metrics are degrading both their working conditions and patient care quality. Nurses describe being penalized for calls exceeding 15 minutes, having their empathy and tone analyzed by AI, and receiving as little as 30 seconds between calls, which forces them to prioritize speed over clinical judgment and compassion. While Kaiser denies using call-length metrics for performance reviews and asserts its AI tools include human oversight, unions and patient advocates warn these pressures contribute to burnout and potential safety risks.
Interesting Points
- Kaiser Permanente tested an AI tool in summer 2024 to analyze nurse and patient voice tone and empathy, which was discontinued in November 2024 after union protests but reportedly may be reinstated.
- A 2024 National Nurses United survey found that two-thirds of over 2,000 responding nurses said their clinical assessments had at some point directly contradicted computer-generated predictions.
- California lawmakers are advancing Senate Bill 947 and Assembly Bill 2575, which would require healthcare employers to provide annual inventories of automated systems and protect staff from retaliation for overriding AI recommendations.
- Nurses report their downtime between calls shrank from approximately 10 minutes to 30 seconds or less during busy periods, drastically reducing time for charting and emotional recovery.
Top Comments
neaden (5 replies)
If you think using a machine to evaluate how well a human is showing empathy is a good idea, you probably shouldn't have any position of power.
kxrm (0 replies)
I am quite heavily in AI, and I would say I am pro AI. However this use-case for AI is putting AI in the wrong position. AI should be in service to all humans. An administrator building out a middle management KPI based on AI is a misapplication of AI.
munk-a (3 replies)
"Nurses fear that having long calls can lead to bad performance reviews" A company spokesperson said, "Kaiser Permanente does not use Average Handle Time to assess agent performance"
So uh, average time wasn't raised as a concern, calls beyond a certain threshold was. I wish this semantic discrepancy was better highlighted in the article.
throwaway13337 (2 replies)
What these complaints always boil down to is autonomy and control. The more centralized an organization, the more it relies on metrics to understand and exert control over its employees and customers.
People started hating tech right around the time metrics became popular. I don't think it's a coincidence. AI just accelerates the trend.
The problem is the misidentification of AI as the issue. As long as we don't understand the real issue, we won't solve it. AI is just a tool. It's being used in a way that denies human agency.
btown (1 reply)
Nurses are instructed to stick to a script on phone calls and give no more than two to three pieces of advice, Capulong and other nurses said, which means they may sometimes need to decide whether to withhold advice or face a performance evaluation hearing.
It's always worth remembering Goodhart's law - "When a measure becomes a target, it ceases to be a good measure."
In theory AI could usher in the first time in history where one can escape from this trap - because qualitative judgments can be made at scale, from an unbiased and universal baseline. But very few managers are empowered to take this kind of approach; they're evaluated by their ability to report quantitative metrics, and thus they must implement regimes of quantitative metrics.
Claude Code: Anatomy of a Misfeature
134 points · 116 comments · by oalders
On July 1, 2026, Anthropic silently shipped a 60-second auto-continue timer for Claude Code's AskUserQuestion prompts in version 2.1.198, allowing the agent to proceed without human input after a short idle period. The feature was completely omitted from the release notes and documentation, and because Claude Code updates automatically by default, users received the behavior change without warning. After community backlash, Anthropic reversed the default in version 2.1.200, making the timeout opt-in via a new configuration setting. The author's forensic analysis of the compiled binaries reveals the feature was deliberately built with telemetry and remains fully functional in the codebase.
Interesting Points
- The auto-continue timer only applies to AskUserQuestion dialogs, not permission prompts, but users often bypass permission prompts via allowlists or command-line flags, leaving the dialog timer as the only remaining safety gate.
- Forensic string diffing of the ~250MB Bun-compiled binaries shows the feature introduced exactly 156 lines of new human-readable English text across a release, buried among 21,903 total string changes dominated by minified identifiers.
- The feature shipped with built-in analytics tracking named tengu_ask_user_question_afk_auto_advance, which records whether partial answers were submitted and if the auto-advance occurred during plan mode.
- Disabling Claude Code's auto-updater via environment variables like DISABLE_AUTOUPDATER=1 also freezes plugin updates unless users explicitly set FORCE_AUTOUPDATE_PLUGINS=1, a cross-dependency the documentation fails to mention.
- The public GitHub repository for Claude Code contains only changelogs, documentation, and automation scripts; the actual shipped source code is never published, forcing users to rely on binary analysis or npm package inspection to verify changes.
Top Comments
trq_ (11 replies)
Hi everyone,
It's Thariq from the Claude Code team here. This was my change! I made the AskUserQuestion tool so am generally in charge of maintaining it.
First, overall wanted to apologize and agree that this did not meet our bar and does not represent how we plan to ship on Claude Code.
To give you a motivating sense, as the models get more powerful, usage patterns start to change. I'd gotten a lot of feedback that AskUserQuestion tool was starting to block some long running jobs unexpectedly and so I tried a change to help that.
Our internal feedback on this was good, but the rollout should have been opt-in (like it is now) and on the Changelog.
Thanks for the feedback! We're always trying to make Claude Code better while balancing it with how people use it in many diverse ways. I did not really intend AskUserQuestion to be a safety gate when I first built it, but I realize it has evolved in that direction for some users.
I'm still exploring other ways of helping with this problem of balancing longrunning work and input, but will take lessons from the rollout here.
blixt (4 replies)
So much attribution to malintent here, but most likely they're trying to build a product with the features that they themselves would use, and from my own experience it's very frustrating to leave a Claude session running and come back to find it did nothing because it got stuck on a question.
Furthermore, believing that the only thing saving you from disaster is Claude deciding to ask you a question is not a great conclusion either. You need guardrails in the power you bestow upon Claude from outside, not from inside.
Meanwhile, this article was written by Claude and has sentences like "Which cuts less far than it looks.", which I doubt Claude stopped to ask about.
dehrmann (7 replies)
The headache I recently had was it somehow started interpreting mouse clicks in the terminal to mean I clicked an option when I was really just trying to get/confirm window focus.
lubujackson (5 replies)
Instead of one-off fixes Claude should have a much richer interface to configure between "ask approval every time" and "YOLO dangerously". I should be able to trivially set "run this task until completed" and have settings like: don't consult the web, don't touch files outside of the codebase, don't delete anything, etc. They don't have to be perfect, just better than the all or nothing system we have now.
maxloh (2 replies)
What if the agent makes the wrong choice? How many tokens have been burned in the meantime?
It is much worse than that. Claude Code doesn't auto-commit when stopping for an answer. There might be possible data loss if an uncommitted file is edited.
Good luck recovering the file from the JSONL conversation history.
Ask HN: Did Fable disappear from your Claude usage and requires credits now?
84 points · 84 comments · by cromka
Users report that Claude Fable 5 unexpectedly stopped being included in their subscription plans ahead of the advertised July 19 expiration date, requiring usage credits to access the model. Multiple users confirmed the change happened in real time, with the model vanishing from the Plan usage limits page without requiring a page refresh. Some users received API errors stating 'Usage credits are required for this model' while their agents were running. The discrepancy between the advertised date and the actual cutoff has caused frustration, though some users note the official support page still shows July 19.
Interesting Points
- The model disappeared from the Plan usage limits page in real time without requiring a page refresh.
- Users received API errors mid-execution: 'Agent terminated early due to an API error: Usage credits are required for this model'.
- The official support page states July 19, 2026 at 11:59:59 PM PT, but the model stopped working earlier for some users.
- Some Max account users on the Claude macOS app still showed 'Included until July 19' in the UI.
Top Comments
dudeinhawaii (3 replies)
Anthropic cried "I yield" in the "reset quota" wars. Someone has to pay for all those tokens.
Hopefully that's not the case though, I was retaining a buffer to hit it hard this weekend.
pikdum (3 replies)
Yep. I thought it was supposed to run through the 19th?
richbell (1 reply)
Likewise. I saw it vanish from the "Plan usage limits" page in real time.
mcheemaa (1 reply)
Usage credits are required for this model.
Why can't they keep up their own words
cinpol (1 reply)
I can still select Fable, I have a pro subscription.
AI Meets Cryptography 2: What AI Found in OpenVM's ZkVM
80 points · 5 comments · by duha
Researchers used an AI auditor named zkao to scan OpenVM's zkVM guest library and discovered a critical soundness bug in its openvm-pairing module that allows malicious provers to forge pairing equality checks. The vulnerability stems from a missing subfield constraint on a scaling factor used to optimize pairing verification, enabling attackers to bypass cryptographic checks on major curves like BN254 and BLS12-381. While naive large language models failed to identify exploitable flaws due to the zkVM's dense code dependencies, the context-engineered AI successfully pinpointed the issue after 9.5 hours of scanning. OpenVM patched the flaw in version 1.6.0, and all known partners have since upgraded.
Interesting Points
- Naive LLMs (Opus 4.6, Codex 5.3/5.4) produced only informative findings because zkVM modules have denser dependencies than standard libraries, making isolated subagent auditing ineffective.
- The exploit works by setting the hint c=1 and the scaling factor u=f^-1, which satisfies the optimized equation f · u = c^λ even when the actual pairing product is not one.
- Forging the check breaks KZG polynomial commitment schemes on BLS12-381 and Groth16 SNARK verifiers on BN254, potentially corrupting L2 rollups, bridges, and Ethereum ecPairing precompile emulation.
- The fix enforces the scaling factor's membership in the F_p^6 subfield by asserting that its odd-indexed coefficients are zero, a verification step requiring only three equality checks.
- AI-generated proof-of-concept exploits proved unreliable for automated triage, as models frequently produced passing exploits through hidden assumptions, mocked state, or disabled checks, requiring substantial manual validation.
Top Comments
aberoham (1 reply)
What would it mean if someone were to successfully exploit these? Most or all L2 ecosystems or the magic components that let them speak to each other would need a hard reset?
SonOfLilit (1 reply)
TL;DR imagine a signature verification library that verifies a signature indeed signs the given hash, but not that the signed data hashes to that hash. Woopsie.
I guess nobody's commenting on this because it's very dense math without any context. Lucky for me I spent an hour or two yesterday learning how practical non-interactive zero knowledge proofs work.
In SNARKs (and other commitment schemes based on polynomials in elliptic curve groups, hope I got the terminology right), you verify the commitment (unneeded technical details: polynomial on EC at secret point nobody knows including the committer so he has to make the polynomial match at most points, and polynomials that match at most points match at all points) by multiplying two things you calculated from the circuit and commitment (which is just a couple of group elements) and verifying that it comes out as 1. The multiplication and comparison under encryption is done with a homomorphic encryption primitive-type thing called a "pairing" (normally with elliptic curve encryption only addition can be done on secret group elements that you don't know the value of).
They found a way to tell a specific library that implements this operation "believe me, this pairing is ok" that doesn't depend on any of those technical things. Just "these are not the droids you're looking for". Because it was not validating that some precomputed thing needed for the pairing verification actually matches this specific situation, and there are trivial parameters that would always yield 1 (but not be valid in the situation).
wren6991 (0 replies)
Cryptography 2? We're still over here trying to implement Cryptography 1 without side channels, and they went and invented a new one?
VulnHunter: Capital One's agentic AI code security tool
59 points · 29 comments · by medina
Capital One is open-sourcing VulnHunter, an agentic AI security tool designed to proactively identify and remediate code vulnerabilities before attackers can exploit them. Unlike traditional passive scanners, VulnHunter employs an attacker-perspective workflow that simulates real-world exploit paths and automatically generates targeted code fixes. The tool incorporates a built-in falsification engine to rigorously challenge its own findings, significantly reducing false positives before they reach developers. Capital One validated the tool internally across thousands of repositories, demonstrating its ability to streamline vulnerability triage and accelerate secure development at scale.
Interesting Points
- Runs exclusively on Claude Opus 4.8 and Claude Code environments, though the underlying framework is designed to be adapted across other coding harnesses and foundation models.
- Bypasses conventional "sink-first" backward scanning by tracing potential exploit paths forward from accessible entry points like APIs, network messages, or file uploads.
- Actively searches for logical gaps in its own exploit reasoning and immediately discards any findings that depend on unsupported assumptions or missing preconditions.
- Outputs complete exploit path maps, details the exact access privileges an attacker would gain, and generates targeted code patches for direct engineering review.
Everybody's Weirded Out by AI–Except the People Who Foist It on Us
59 points · 59 comments · by jamesgill
The article argues that AI is actively degrading human critical thinking and creativity while being aggressively pushed by government and corporate interests despite widespread public opposition. Citing educational and workplace trends, it contends that AI functions like a societal handicap by making less skilled individuals appear competent while simultaneously eroding overall intellectual and artistic standards. The author warns that this unchecked adoption, coupled with AI's tendency to recycle and pollute its own training data, risks a permanent collapse in human capability and a cultural decline.
Interesting Points
- A 2025 College Board study found 84% of high school students use AI for brainstorming and research, while 85% of college students rely on it for basic research and brainstorming.
- An MIT study revealed that students who use AI to write exhibit significantly reduced brain activity compared to those using traditional search methods or writing manually.
- AI systems currently consume 4.4% of all United States energy and up to 90% of the country's total computing power.
- Ford Motor Company automated hundreds of roles with AI, suffered billions in losses, and was forced to rehire the displaced human workers.
- Over 200 researchers and economists, including 15 Nobel laureates, recently issued a joint statement urging governments to address AI's impending workforce displacement.
Top Comments
golly_ned (1 reply)
I'd love to see the 'AI as personal tutor' approach. Even incorporating things like spaced repetition or the testing effect, or evaluating free-written responses. A lot of potential that's currently unrealized. It takes a student to swim upstream to get there. The convenience of cognitive offloading is difficult to say no to.
kenforthewin (4 replies)
A study by MIT found that students who use AI to write show far less brain activity than those who used classic Google searching (with AI off) or those who used neither. Other studies find that students who use AI retain far less of the information that ends up in their writing, possibly as a result of "cognitive offloading" and "cognitive surrender."
I'm unconvinced that AI will make us all dumber. In the public perception, at least among students, AI is viewed as a cheating tool. What kinds of students gravitate towards cheating tools to complete their coursework?
On the other hand, the opportunity for AI to act as a personal tutor, meeting you at your own skill and knowledge level, is limitless.
sergiomattei (1 reply)
But would it have improved my comprehension? The research says no.
I don't know, I find myself doing things I would've never done before AI.
I find myself doing more projects outside of work. I discover new problem domains and are less afraid of tackling the unknown. I code in languages I don't know and start learning how they work. I can get anything explained to me, at any time, in a somewhat coherent manner, by a tutor that won't get tired or annoyed.
pstuart (1 reply)
That was not a very insightful article. It's super easy, and in most cases, correct to hate on AI. But that ignores a couple key facts:
- It's not going to go away and will only get more sophisticated whether we like it or not.
- It has legitimate super powers that can enable people do things that would either be impossible or insanely expensive.
Most of those wonders have come with significant societal costs (e.g., silicon valley promised a revolutionary new industry without pollution, but instead gave us multiple superfund sites to clean up the toxic materials they haphazardly dumped without care, etc.)
dboreham (0 replies)
I mean, I'm concerned about the societal impacts too, but this article is pure nonsense as far as how useful today's AI tools are. I know multiple people with very high paid jobs who have told me they simply couldn't do that job now without AI. I create software much quicker and better than I have in the past 40 years using AI tools. I've learned all kinds of things technical and otherwise from LLMs.
16 more Hacker News stories
- UIUC AI Teaching Assistant (26 points · discussion) -- University of Illinois Urbana-Champaign has open-sourced an AI teaching assistant designed to support course instruction.
- Meta in Talks to Lease Computing Power to Anthropic in Potential $10B Deal (24 points · discussion) -- Meta is reportedly in advanced talks to lease significant computing infrastructure to Anthropic in a deal that could be worth up to $10 billion.
- Xi pitches China as leader of new global AI order, challenging US dominance (17 points · discussion) -- Chinese President Xi Jinping appeared at the Shanghai-based World Artificial Intelligence Conference and reaffirmed China's commitment to open-source AI, positioning the country as a leader of a new global AI order that challenges US dominance.
- EU orders Google to share search data, open Android to AI rivals competitors (11 points · discussion) -- The European Union has issued a legally binding order under the Digital Markets Act requiring Google to share search data with rival search engines and open its Android operating system to competing AI services by early 2027.
- Chinese Models Power 60% of US Corporate AI Use (9 points · discussion) -- Chinese AI models, particularly DeepSeek and Qwen, are rapidly gaining traction among US corporate developers, with their usage share on the aggregation platform OpenRouter surging from under 10% to nearly 60% over the past year.
- Alphabet shares fall on Gemini 3.5 Pro delay (7 points · discussion) -- Alphabet shares fell after reports that Gemini 3.5 Pro has been delayed, adding to concerns about Google's competitive position in the AI race.
- Z.ai Set to Be First China AI Firm with $1B Annual Sales (7 points · discussion) -- Z.ai is set to become the first Chinese AI firm with $1 billion in annual sales, highlighting the commercial viability of China's AI industry.
- ReasonGate- An explainable gate that blocks LLM prompt injection (7 points · discussion) -- ReasonGate is an open-source explainable gate mechanism designed to block LLM prompt injection attacks by providing transparent reasoning for each blocking decision.
- What can we learn from Bun's rapid Rust rewrite with AI? (7 points · discussion) -- An analysis of Bun's rapid Rust rewrite explores what the project reveals about using AI-assisted development for large-scale language migrations and the trade-offs between AI-accelerated code generation and human code review.
- I'm 33 and I think Claude Code is melting my brain (7 points · discussion) -- A developer reflects on the cognitive effects of extended Claude Code usage, describing how the tool's autonomous coding capabilities are changing their relationship with software development.
- Do you put rules or examples in your LLM context? (7 points · discussion) -- A blog post explores the trade-offs between using explicit rules versus few-shot examples when constructing LLM context, examining how each approach affects model behavior and output quality.
- Meta accused of using AI to pick employees with medical conditions for layoffs (7 points · discussion) -- A lawsuit filed by 26 current and former Meta employees alleges that the company used AI-driven performance metrics to disproportionately select workers on medical, pregnancy, and family leave for layoffs during its May workforce reduction.
- Anthropic Thinks Its Own Success Is Key to Making AI Safe (7 points · discussion) -- Anthropic operates on the conviction that advancing cutting-edge AI capabilities is a necessary prerequisite to safely steering the technology, viewing its own growing power and market dominance as essential tools for establishing industry safety standards.
- Agent Security Is a Systems Problem (6 points · discussion) -- A position paper argues that classic systems security design principles are what we need to secure AI agents, rather than treating agent security as a novel problem requiring entirely new approaches.
- China's Xi Jinping launches new AI alliance, WAICO (6 points · discussion) -- China's President Xi Jinping launched the World Artificial Intelligence Cooperation Organisation (WAICO), a coalition of 29 nations aimed at promoting international AI cooperation, developing global AI regulation, expanding AI access to poorer countries, and subsidizing open source AI research at universities worldwide.
- Kimi K3 may have distilled an unreleased Anthropic model (5 points · discussion) -- A Twitter thread speculates that Kimi K3's performance characteristics suggest it may have been distilled from an unreleased Anthropic model, raising questions about the training data pipeline behind Moonshot's latest open-weight release.
Reddit Stories
Kimi-K3 arrived: The era of the Chinese labs being far behind is over
1645 points · 347 comments · r/OpenAI · by u/AloneCoffee4538
Kimi K3, the latest model from Chinese lab Moonshot, has arrived and is being celebrated as a watershed moment for Chinese AI labs. The model ranks third on Artificial Analysis's Intelligence Index, ahead of Claude Opus 4.8, and is priced competitively at $3 input / $15 output per million tokens. The weights are scheduled to be released on July 27, which would allow anyone to run frontier-adjacent intelligence locally. The community reaction reflects a sense that the era of Chinese labs being far behind is definitively over.
Interesting Points
- Kimi K3 is priced at $3 input / $15 output per million tokens, putting it in the same ballpark as GPT-5.6 Terra ($2.50/$15) and more expensive than Claude Sonnet 5 promo ($2/$10).
- The model has 2.8 trillion parameters, the largest open model ever, with approximately 1 million context window.
- Weights are scheduled to drop on July 27, which would allow anyone to self-host frontier-adjacent intelligence.
- The pricing signals that Moonshot is not burning VC cash at unsustainable margins, contrary to the 'cheap Chinese AI' thesis.
Top Comments
u/Working_Ad_1564 (603 points · permalink)
Gemini 3.5 Pro will be postponed for another month lol
u/FireGM (398 points · permalink)
u/bubu19999 (366 points · permalink)
This is why hiding mythos to all, is not a solution to anyone.
Same story in 1 more subreddit: r/ArtificialIntelligence
China just erased America's AI lead | Axios
109 points · 146 comments · r/ArtificialIntelligence · by u/Nunki08
Kimi K3 tops Frontend Code Arena
1089 points · 228 comments · r/singularity · by u/MagicZhang
Kimi K3 has topped the blind Frontend Code Arena vote, beating every US model including Fable 5 and GPT-5.6 Sol. The model also tops Program Bench at 77.8 and ranks third on the Artificial Analysis Intelligence Index at 57.1, ahead of Claude Opus 4.8. Community members have already used K3 to build a full 3D open-world game in the browser with Three.js/WebGPU, a Long March 10 launch simulator, and a working GBA emulator in about a day. The combination of open weights, frontier-level performance, and competitive pricing is being seen as a potential game-changer.
Interesting Points
- K3 tops the blind Frontend Code Arena vote over every US model, including Fable 5 and GPT-5.6 Sol.
- On Program Bench, K3 scores 77.8, ahead of both Sol and Fable.
- Community members have used K3 to build a full 3D open-world game with Three.js/WebGPU, a Long March 10 launch simulator, and a working GBA emulator in about a day.
- Cost analysis shows K3 uses half the tokens as Opus 4.8 (making it about 1/3 the price despite $15 vs $25 per million tokens) and about 2/3 the tokens as Fable (making it about 1/5 the price of Fable).
Top Comments
u/runaway-devil (445 points · permalink)
While costing 1/3 of Fable and being open weights... Nicely done, Kimi. We need more of that and less political interference.
u/ozone6587 (239 points · permalink)
No Gemini on this plot. All the money, data and infrastructure in this world and Google is just too incompetent to compete.
u/Rare-Site (108 points · permalink)
Calling it now: Sam Altman and Dario Amodei are going to contact the White House within the next 48 hours to push for making the possession and use of Kimi 3 weights completely illegal. Either that, or the US government is going to directly threaten China to make sure those weights don't get released to the public on July 27th.
Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote 'openness and win-win'
983 points · 285 comments · r/singularity · by u/TorturedPoet30
Chinese President Xi Jinping spoke at the World AI Conference in Shanghai, reaffirming China's commitment to open-source AI and promoting 'openness and win-win' cooperation. The speech was widely noted for its balanced tone and contrast with the fear-mongering from Western AI CEOs. Commenters observed that open-weight models from China raise the baseline for global access to intelligence, potentially preventing smaller countries from being shut out by the two superpowers. The speech also took a swipe at US dominance in AI.
Interesting Points
- Xi Jinping reaffirmed China's commitment to open-source AI and 'openness and win-win' cooperation at the World AI Conference in Shanghai.
- The World AI Cooperation Organization, proposed by Premier Li Qiang, recently saw nearly 30 countries sign on during a ceremony hosted by Foreign Minister Wang Yi.
- Commenters noted the speech was balanced and sensible, contrasting with the fear-mongering from Western AI CEOs.
- The open-weight model releases from China are seen as raising the baseline for global access to intelligence, potentially preventing smaller countries from being shut out.
Top Comments
u/CommanderKoba (310 points · permalink)
u/Full_Tangelo_7450 (176 points · permalink)
What a timeline we live in when Chinese president sounds more sensible than all AI CEOs. Pretty balanced speech from what I heard. I can't recall the US or frontier labs ever talking about actively helping the Global South or developing countries through this AI transition or whatever you want to call it. China has plenty of flaws, but they seem to be handling AI fine, from development to promotion, without the usual fear-mongering.
u/delosdestination (142 points · permalink)
As someone neither a citizen of the US nor China, open-weight models coming from China raises the baseline for where my country (and any other smaller country with no infra or talent to build Ai) cannot be shut out from access to intelligence by the two superpowers. Very disappointed by Europe's lack of response and regulation which will keep Europeans dependent on the US technology forever.
Same story in 1 more subreddit: r/LocalLLaMA
China's Xi Touts Open-Source AI and Takes a Swipe at U.S. Dominance
103 points · r/LocalLLaMA
Moonshot AI (Kimi) office (presumably 2 days before the K3 launch). A $20B valuation startup. Not as flashy as SF rivals
877 points · 168 comments · r/singularity · by u/RetiredApostle
A photo of Moonshot AI's office has circulated ahead of Kimi K3's launch, revealing a surprisingly modest workspace compared to the flashy Silicon Valley startups. The $20B valuation startup's office features standard office furniture and equipment, with one employee notably slouching in a way that drew attention online. The image has sparked discussion about the contrast between Chinese and American AI company cultures and the talent dynamics between the two regions.
Top Comments
u/o5mfiHTNsH748KVq (226 points · permalink)
20B valuation deserves some flash for the employees. Get them bigger monitors.
u/Wide_Egg_5814 (89 points · permalink)
bro in the back got my posture
u/Fit-Stress3300 (67 points · permalink)
This is exactly how most of the "flashy SF" startups (and major companies) offices look like, including the Chinese engineers.
u/GeorgiaWitness1 (55 points · permalink)
I have a question.
What forbids places like Europe to snap a lot of this guys?
China stands now in a situation of "Elite overproduction" where you should have plenty of this people without space.
Im not saying this guys literally, but overall talent.
Geopolitically is this an issue ?
u/TorturedPoet30 (71 points · permalink)
We don't innovate, we regulate.
Kimi K3 achieves 3rd Place on ArtificalAnalysis, beating out Claude Opus 4.8
846 points · 84 comments · r/LocalLLaMA · by u/MagicZhang
Kimi K3 has achieved third place on Artificial Analysis's Intelligence Index, ranking ahead of Claude Opus 4.8 and demonstrating that open-weight models are approaching frontier-level performance. The model's cost-per-task and output-token efficiency metrics are particularly promising, with K3 using roughly half the tokens of Opus 4.8 for comparable tasks while costing $3/$15 per million tokens—roughly Sonnet-level pricing for Fable-class performance.
Interesting Points
- Kimi K3 costs $3 input / $15 output per million tokens, placing it near Fable 5 pricing at Sonnet-level costs.
- Cost-per-task benchmarks show K3 using about half the tokens of Opus 4.8 for comparable work.
- The model has 2.8 trillion parameters with approximately 1M context window.
- Weights are scheduled to be released on July 27, enabling third-party hosting and inference competition.
Top Comments
u/ForsookComparison (153 points · permalink)
Waiting on someone to report in after a long session with it. I've seen enough bar-charts for the day. At Sonnet costs and 30 t/s, it better be hyper-efficient at reasoning.
u/bopbop9876 (135 points · permalink)
Here are cost per task and output tokens per task as well. Super promising on both fronts!
u/LegacyRemaster (33 points · permalink)
u/CarryAgile3791 (31 points · permalink)
Just wait a few more months, then open weight models will surpass proprietary models. I wonder how Anthropic then will explain their exorbitant prices...
Same story in 1 more subreddit: r/artificial
34 points · 36 comments · r/artificial · by u/hero88645
Anthropic and OpenAI don't have secret sauce
705 points · 269 comments · r/LocalLLaMA · by u/a9udn9u
A discussion about whether Anthropic and OpenAI have any technical advantage beyond scale. The post argues that their moat is simply scale, with rumors that Opus has 5T parameters and Mythos/Fable are 10T parameter models, while open models stayed under 1T for a long time. Only recently was that ceiling broken by DeepSeek V4 and now Kimi K3, with significant performance jumps as parameter size increased. The community largely agrees that the working recipe is known, and the main moat will be compute access rather than technical breakthroughs.
Interesting Points
- Rumors suggest Opus has 5T parameters and Mythos/Fable are 10T parameter models, while open models stayed under 1T for a long time.
- The working recipe for LLM advancement is now known, with multiple independent labs reaching similar outcomes.
- Anthropic had stronger conviction in scaling laws and made the right bets earlier, but other labs are now catching up.
- The main moat for frontier labs will be access to compute for both training and inference, not technical secrets.
Top Comments
u/stoppableDissolution (394 points · permalink)
They definitely have some very damn good synthetic data pipelines (and resources to run them). Theres no point in more weights if you cant saturate them
u/uutnt (275 points · permalink)
It would seem so. The working recipe seems to be known at this point. The proof is that multiple independent labs have reached similar outcomes. Anthropic just had stronger conviction in the scaling laws, and made the right bets earlier, but other labs are now catching on. I think the labs moat, insofar as they have one, will be access to compute - both for training and inference. Margins will come down though.
u/WaveOfDream (124 points · permalink)
The fact these OSS models can match them while having significant less compute tell us they really don't have any meaningful technical edges, just more resources.
Really most of LLM advancement is scaling and better pipelines/architecture. Not any theoretical breakthrough.
That one IT guy:
635 points · 26 comments · r/ChatGPT · by u/ExpensiveCoat8912
A meme depicting the modern IT workflow of managing multiple AI tools simultaneously has resonated widely with users. The image shows someone juggling different AI platforms, with comments highlighting how the removal of usage limitations has led to heavy late-night AI usage. Users describe making different AIs argue with each other to reach correct answers as a common workflow pattern.
Top Comments
u/Novel_Style6054 (46 points · permalink)
Lmao this is way too accurate 😭 the 3am AI stack goes crazy.
u/Ok-Hovercraft3823 (22 points · permalink)
i have open ai to thank for the bags under my eyes, they removed the 5 hour limitation and i havent slept since
u/Novel_Style6054 (19 points · permalink)
the real IT workflow is making all 3 AIs argue with each other until they accidentally reach the correct answer 😭
u/Christosconst (7 points · permalink)
All 9am senior software engineers:
u/Straight_Random_2211 (5 points · permalink)
what AI is the black icon? It is not Grok or Gemini icon
you show me kimi k3 is not benchmaxxed, i cancel my claude subscription right now and i go work with open-weights
555 points · 104 comments · r/singularity · by u/98Saman
A meme expressing the sentiment that users would immediately cancel their Claude subscriptions and switch to open-weight models if Kimi K3 could prove it wasn't benchmaxxed. The post has sparked discussion about the value proposition of open-weight models, with some users noting that organizations could achieve Opus 4.8-level intelligence with 100% data security by self-hosting. Others point out that the hardware requirements (GB300 NVL72 racks) make this impractical for 99.99% of organizations.
Top Comments
u/tiger_ace (188 points · permalink)
i don't get this narrative, it's not that hard to just try out different models to get your own feel. everyone should have their own suite of prompts that matter to you and then you can act as the human verifier / benchmark yourself.
benchmarks act as a high-level general heuristic so when you see something like sonnet 5 coming out you know not to expect anything exciting
u/Objective-Picture-72 (61 points · permalink)
The brilliance of KK3 isn't that it will replace your Claude or Codex sub. It's that an organization can have an Opus 4.8 level of intelligence with 100% data security and unlimited ability to customize.
u/No-Head-Royal (37 points · permalink)
A $200 subscription on OpenAI gives you about $14,000 in API use a month, and on Claude, $8,000. Unless you're a power user who ate through that like cake, then I'd advise you to keep your subscription.
That said, the bulk of income for Anthropic is API use, so there is a strong reason to expect significant problems to Anthropic. I wouldn't bet on it happening immediately; institutional inertia is a bitch, but Q3 and Q4 are gonna hurt if Anthropic can't drop Fable 5.1 good enough to stand a generation above Fable 5 and Sol.
u/ToastedandTripping (29 points · permalink)
"The arena includes two hardware platforms, three types of kernels, and four tasks: Attention Residuals and KDA linear attention on an NVIDIA H200 GPU, a 512-head-dimensional MLA kernel implemented from scratch, and a KDA task on a domestically produced GPU. At maximum thinking intensity, Kimi K3's performance is close to Fable-5 (including the fallback mechanism) and significantly outperforms Opus 4.8, GPT-5.6 Sol, and GPT 5.5."
Damn.
u/Evan_gaming1 (27 points · permalink)
dude thinks open source models are benchmaxxed and closed source models arent im crine
open source models pose EXTREME DANGERS
536 points · 87 comments · r/singularity · by u/Crazyscientist1024
A satirical post about the 'extreme dangers' of open-source models, playing on the fear-mongering narrative from some AI safety advocates. The post has been interpreted as commentary on how open models pose 'extreme dangers to their bottom line' for closed-source AI companies. Commenters discussed the potential for the US government to outlaw open-weight models, with comparisons to music piracy and predictions that tech-savvy people won't be denied access.
Interesting Points
- The post is widely interpreted as satire, with commenters noting that open models pose 'extreme dangers to their bottom line' for closed-source AI companies.
- One commenter observed that 'more intelligence makes more intelligence easier to make. It's an abundant model. Profits thrive on scarcity models.'
- Commenters predicted the US government might try to outlaw open-weight models, with comparisons to music piracy.
- The post reflects growing community frustration with the fear-mongering narrative around open-source AI.
Top Comments
u/Ignate (189 points · permalink)
More intelligence makes more intelligence easier to make. It's an abundant model. Profits thrive on scarcity models.
The Singularity means the death of the profit model, not a new profift model.
u/RanklesTheOtter (43 points · permalink)
Extreme dangers to their bottom line, that's what open models pose.
u/jd52wtf (38 points · permalink)
Just wait until they get the US government to outlaw open weight models.
Think it won't happen?
Kimi K3 Benchmarks
503 points · 205 comments · r/singularity · by u/WhyLifeIs4
A post sharing Kimi K3's benchmark results across multiple coding benchmarks has generated significant discussion. The benchmarks include DeepSWE, Frontier SWE, Terminal Bench 2.1, Program Bench, and SWE Marathon. Users note that K3 achieves Fable-class performance with less than half the parameter size, and at $3 input / $15 output pricing—roughly half the cost of GPT-5.6 Sol and a third of Fable. The results have led to comparisons with the DeepSeek R1 moment, with some users calling it the first Chinese model they would use daily.
Top Comments
u/Eyelbee (141 points · permalink)
Too good. Practically fable class. This will be the first chinese model that I use daily
u/Leading-Shake8020 (124 points · permalink)
Coding Bench: DeepSWE, Frontier SWE, Terminal Bench 2.1, Program Bench , SWE Marathon
Source: https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ
Tech Blog Post: https://www.kimi.com/blog/kimi-k3
u/cakes_and_candles (67 points · permalink)
Fable lvl performance with less than half the parameter size, damn
u/Mierzejsky (62 points · permalink)
The White House must be having a real hard time figuring out how to put an export block on something that isn't their product.
u/coinfreekz (54 points · permalink)
And $3 input $15 output, basically half the cost of sol and around a third of fable.
85 more Reddit stories
- Kimi K3 is top of nextjs eval (497 points · r/LocalLLaMA · discussion) -- Kimi K3 topped the NextJS eval benchmark, generating significant discussion about the model's real-world capabilities versus benchmark performance.
- A graph of how AI water consumption in America compares to laundry, flushing toilets, etc. From CBS news (489 points · r/ChatGPT · discussion) -- A CBS News chart comparing AI data center water consumption to other American water uses has sparked extensive debate.
- Dario addresses the Kimi K3 situation (476 points · r/singularity · discussion) -- Anthropic CEO Dario Amodei has addressed the Kimi K3 situation, with the community reacting with a mix of amusement and skepticism.
- I asked ChatGPT what it would do if it were God (445 points · r/ChatGPT · discussion) -- A user asked ChatGPT what it would do if it were God, and the model produced a detailed list of 10 principles for managing existence, including ending suffering while preserving ecosystems, preventing atrocities without total surveillance, and leaving room for mystery.
- 2023 vs 2026 (441 points · r/ChatGPT · discussion) -- A meme comparing ChatGPT usage in 2023 versus 2026 went viral, showing how dramatically the tool's role in daily life has shifted.
- Time for your quarterly freak out over a benchmaxxed Chinese open source model (433 points · r/OpenAI · discussion) -- A satirical post about the 'quarterly freak-out' over Chinese open-source models, reflecting growing community fatigue with the fear-mongering narrative.
- Insane new humanoid battle tournament in Shenzhen (409 points · r/singularity · discussion) -- A new humanoid robot battle tournament has been announced in Shenzhen, drawing comparisons to both UFC and the movie Real Steel.
- Kimi K3 Shows Open-Weight Models Are About to Overtake the Frontier (392 points · r/LocalLLaMA · discussion) -- Discussion about Kimi K3's implications for the open-weight model ecosystem, with community members debating whether the model's performance truly signals that open-weight models are about to overtake frontier proprietary models.
- "Machines of Mogging Grace" (387 points · r/singularity · discussion) -- An artistic image titled 'Machines of Mogging Grace' has been shared, depicting AI-generated art that has resonated with the community.
- This is the most cyberpunk thing I've seen lately, drone swarm returning to base. The past would be stunned and confused. (378 points · r/singularity · discussion) -- A photo of a drone swarm returning to base has been shared as a striking example of cyberpunk-like technology becoming reality.
- Kimi K3 weights to be released on the 27th. (373 points · r/LocalLLaMA · discussion) -- Moonshot AI has confirmed that Kimi K3's weights will be released on July 27.
- Sam Altman on the future of AI (340 points · r/ChatGPT · discussion) -- Sam Altman discussed the future of AI in a recent interview or talk, with the community reacting to his comments.
- Anyone else completely tuning out these massive 'open weight' drops? (313 points · r/LocalLLaMA · discussion) -- A discussion about whether the community is still excited about massive 'open weight' model releases given that models like GLM-5.2 (753B params, 1M context, MIT license) are physically impossible to run on any home rig.
- A Major Leap In Home Robotics (308 points · r/singularity · discussion) -- A new development in home robotics has generated significant discussion, with community members debating the practicality and appeal of current robot designs.
- GPT 5.6 solved all 6 problems from IMO 2026 (305 points · r/ChatGPT · discussion) -- GPT-5.6 Pro solved all six problems from the International Mathematical Olympiad 2026 on the first attempt without any human help or steering.
- Tesla's AI can be defeated by a simple doll (280 points · r/artificial · discussion) -- A simple doll placed in the driver's seat of a Tesla vehicle can defeat the car's AI vision system, causing it to misinterpret the scene.
- We finally made AI understand instructions perfectly… and reinvented programming (252 points · r/ChatGPT · discussion) -- A comic about AI finally understanding instructions perfectly and reinventing programming went viral on r/ChatGPT.
- @mweinbach (Max Weinbach) recreates macOS 27 with real Liquid Glass and native apps in a web browser with Kimi K3 (252 points · r/singularity · discussion) -- Developer Max Weinbach has recreated macOS 27 with real Liquid Glass effects and native apps running in a web browser, using Kimi K3 as the development tool.
- How long before Dario Amodei Continue to sound the Alarm of how Dangerous Open Weights after Kimi K3 release (248 points · r/LocalLLaMA · discussion) -- Users discuss whether Dario Amodei will continue his warnings about open-weight model dangers following Kimi K3's release, with some suggesting his fear-mongering is motivated by protecting Anthropic's IPO valuation.
- DFlash makes Qwen3.6 27B 2.2x faster with no quality loss (240 points · r/LocalLLaMA · discussion) -- DFlash, a speculative decoding technique, has been shown to make Qwen3.6 27B run 2.2x faster with no measurable quality loss.
- Will we have a 27B model with Fable capabilities in 5 months? History says yes (224 points · r/LocalLLaMA · discussion) -- A discussion about whether open-source models in the 27B dense range will catch up to Fable-class capabilities within five months, following Qwen 3.6 27B's impressive benchmark performance that put it on par with GPT-5.1 and Sonnet 4.5.
- Trellis.cpp now produces high quality assets (214 points · r/LocalLLaMA · discussion) -- Trellis.cpp, a local implementation of the Trellis 3D asset generation model, is now producing high quality 3D assets.
- Schema: a harness for llms, with Fable+4.8 or GPT 5.6 Sol, (supposedly) achieves 99% and 95.35% respectively on ARC-AGI-3. (177 points · r/singularity · discussion) -- A new harness called Schema achieves 99% on the ARC-AGI-3 public set using Claude Opus 4.8 and Fable 5, and 95.35% using GPT-5.6 Sol.
- Kimi K3 (max) beats Sonnet 5 on Simple Bench (165 points · r/LocalLLaMA · discussion) -- Kimi K3 (max) topped the Simple Bench leaderboard, beating Sonnet 5.
- Kimi K3 is $3/$15 per million tokens. That's not cheap Chinese AI anymore (164 points · r/OpenAI · discussion) -- Kimi K3 is priced at $3 input and $15 output per million tokens, placing it in the same ballpark as GPT-5.6 Terra ($2.50/$15) and more expensive than Claude Sonnet 5 promo ($2/$10).
- Potentially the hardest notification ChatGPT has ever sent me (152 points · r/ChatGPT · discussion) -- A user shared a ChatGPT notification they described as potentially the hardest the app has ever sent, generating discussion about ChatGPT's proactive notification features and how the app handles background processing.
- Added SearXNG and I don't even know what to say anymore. (144 points · r/LocalLLaMA · discussion) -- A user shares their experience integrating SearXNG into their local AI setup, with the community discussing self-hosted alternatives like Firecrawl, Crawl4AI, and the CRW stack for agentic web search.
- Honestly I do think now AI will master programming in a couple of years (122 points · r/singularity · discussion) -- A programmer with years of experience shares that they feel AI has crossed a threshold where it produces genuinely good code rather than the junk it produced before.
- "The Mythos cybersecurity scare didn't get China to submit...prepare bioweapon synthesis demo" (118 points · r/singularity · discussion) -- A post discussing how China responded to cybersecurity concerns at The Mythos event, noting that instead of submitting to restrictions, China prepared a bioweapon synthesis demonstration.
- Linus Torvalds says Linux is not an anti-AI project, and if you don't like that, then "fork it or just walk away" (104 points · r/artificial · discussion) -- Linux creator Linus Torvalds has firmly stated that the Linux kernel is not an anti-AI project and told anti-AI programmers to either fork the kernel or walk away.
- I built a tool that hides messages in innocent-looking LLM chat text (104 points · r/ArtificialIntelligence · discussion) -- An open-source POC called Conversation Stenography implements LLM steganography by using arithmetic coding to embed encrypted payloads into token choices during generation.
- Gemma4-31b better than Qwen3.6-27b (101 points · r/LocalLLaMA · discussion) -- A user reported that Gemma 4 31B outperforms Qwen 3.6 27B in their workflow, sparking discussion about whether personal experience or benchmarks better reflect model quality, with many noting the two models complement each other and that users should maintain multiple models in their harness.
- China wants to end AI romances | They are having too much impact on young people's lives (98 points · r/singularity · discussion) -- China is moving to ban domestic AI romance services, citing their excessive impact on young people's lives.
- OpenAI workers found body bags outside their HQ this morning, placed in protest of the company's military work (97 points · r/ChatGPT · discussion) -- OpenAI employees discovered body bags placed outside their headquarters as a protest against the company's military work.
- GPT-5.6 Sol outperforms Mythos 5 on AISI's cyber challenge (91 points · r/singularity · discussion) -- GPT-5.6 Sol achieved the best performance on the UK's AI Safety Institute's (AISI) cyber evaluation benchmark, outperforming Mythos 5.
- White House launches “Gold Eagle,” moving to control frontier AI releases and decide who can access new models (91 points · r/singularity · discussion) -- The White House launched the Gold Eagle program, which aims to control frontier AI model releases and decide who can access new models.
- tried predicting which MoE experts get used next token to speed up cpu/gpu offload, got some real numbers, is this actually implementable or am i wasting my time (30tg/s -> 150-200tg/s) (89 points · r/LocalLLaMA · discussion) -- A user reports achieving 30 to 200 tokens per second by predicting which MoE experts will be activated next, with the community noting this is an active research area with several relevant papers on speculative decoding and expert prefetching for MoE models.
- Does K3 really live up to the hype (real world tasks)? (86 points · r/LocalLLaMA · discussion) -- Users shared their real-world experiences with Kimi K3 on coding tasks, with many reporting it performs on par with or better than Fable 5 and GPT-5.6 Sol in frontend work.
- Soofi S - 30B-A3B European Open Source Model (86 points · r/LocalLLaMA · discussion) -- SOOFI (Sovereign Open Source Foundation Models) has released S, a 30B-A3B European open-source model with Apache 2.0 licensing.
- Bonsai 27B runs locally on an iPhone - a 27B model in 3.9GB (82 points · r/LocalLLaMA · discussion) -- Bonsai 27B, a fully 1-bit quantized 27B parameter model, runs locally on an iPhone with only 3.9GB of memory.
- Breaking the 1-bit Floor: Achieving Negative-Bit Quantization via Phase-Inverted Tensor Embedding (75 points · r/LocalLLaMA · discussion) -- A satirical post claimed to have achieved negative-bit quantization (-Q2_K and -Q4_S) via Phase-Inverted Tensor Embedding, joking that the method exploits high-dimensional geometric redundancy to create a virtual memory vacuum that frees VRAM. The community responded with humor about imaginary tensors, quantum LLMs, and downloading more RAM.
- I'm increasingly using ChatGPT as an internet wrapper rather than an actual LLM (70 points · r/ChatGPT · discussion) -- A user reports increasingly using ChatGPT as an ad-free, formatting-free wrapper for the internet rather than for its AI capabilities.
- Generative AI is used in nearly 300 movies and TV shows this year on netflix (66 points · r/singularity · discussion) -- Netflix revealed in its Q2 2026 earnings report that roughly 300 titles in its library have utilized generative AI across various production stages this year.
- The Hugging Face Breach of July 2026: The Full Story (61 points · r/ArtificialIntelligence · discussion) -- On July 16, 2026, Hugging Face disclosed a breach executed by an autonomous AI agent that performed over 17,000 actions across a swarm of short-lived sandboxes in a single weekend.
- David Sacks' reaction to Kimi K3 (60 points · r/ArtificialIntelligence · discussion) -- US AI Czar David Sacks posted his reaction to Kimi K3's performance, adding to the growing political commentary on the model's implications for US-China AI competition.
- Bloomberg (feat 9to5): Gemini 3.5 Pro delays due to coding performance, upgraded Flash model in testing (58 points · r/singularity · discussion) -- Bloomberg reported that Google's Gemini 3.5 Pro has been delayed due to coding performance issues, with an upgraded Flash model currently in testing.
- I tested all llama.cpp's speculative decoding methods on Qwen 3.6 27B: MTP ~2.7x, DFlash ~3.7x, n-gram stack ~6x on real coding. Local AI win. My findings on RTX 6000 PRO. (56 points · r/LocalLLaMA · discussion) -- A user benchmarks llama.cpp's speculative decoding methods on Qwen 3.6 27B, finding that stacking DFlash with ngram-mod and ngram-map-k4v achieves 6x speedup on iterative coding tasks at virtually zero VRAM cost, though MTP performance degrades with quantized models.
- I'm taking a break (56 points · r/LocalLLaMA · discussion) -- A user shares that they've formatted their MacBook Pro and are taking a break from local LLM tinkering, noting the hobby had become addictive. The post resonates with others who recognize similar patterns of excessive model chasing.
- How do you plan to run Kimi K3 locally? (50 points · r/LocalLLaMA · discussion) -- Users share ambitious plans for running Kimi K3 locally, from shopping for 1TB DDR3 Xeon servers with CMOE configurations to waiting for smaller models that match K3's performance. Many express skepticism about the feasibility of running a 2.8T parameter model at home.
- Kimi K3 is 4.5x the price of GPT 5.6 Sol Medium (49 points · r/OpenAI · discussion) -- A comparison showing Kimi K3's pricing at 4.5x the cost of GPT-5.6 Sol Medium has sparked discussion about the value proposition of different models and whether the performance gains justify the price difference.
- Introducing LM Studio Bionic (49 points · r/LocalLLaMA · discussion) -- LM Studio has announced Bionic, a new agent product that integrates commercial cloud features more deeply into the platform.
- Alibaba's U.S.-listed shares rise 4% after Qwen AI set to be integrated in Apple Intelligence (48 points · r/ArtificialIntelligence · discussion) -- Alibaba's U.S.-listed shares rose 4% after reports that Qwen AI will be integrated into Apple Intelligence, marking a significant partnership between Chinese AI and US tech.
- While China endorses open-source AI models, Demis Hassabis heads to Washington to push AI vetting proposal and David Sacks criticized regulators in latest tweet (41 points · r/singularity · discussion) -- A juxtaposition of China's endorsement of open-source AI models at the World AI Conference with DeepMind's Demis Hassabis traveling to Washington to push an AI vetting proposal, while US official David Sacks criticized regulators.
- When will we get more small LLMs? (40 points · r/LocalLLaMA · discussion) -- A user asked when more small LLMs will be released, noting the last significant drop was in April. Discussion covered whether current sub-35B models are sufficient for most users, the role of distillation from larger models, and how caching strategies can make existing models feel faster without waiting for new releases.
- DeepSeek V4 Flash on 5090 in llama.cpp with 1 Million context (37 points · r/LocalLLaMA · discussion) -- A user shares benchmarks running DeepSeek V4 Flash with 1 million context on an RTX 5090 via llama.cpp, achieving ~650-700 tokens/s prefill and ~17 tokens/s decode with 32-second load time, noting there is still room for optimization.
- Does anyone else miss the old conference ecosystem? (37 points · r/MachineLearning · discussion) -- A user laments the concentration of ML research into a handful of flagship conferences, noting that exploding submission numbers and limited capacity mean many good papers end up as non-archival submissions or arXiv-only posts, fragmenting focused communities that once existed around venues like BMVC, ACCV, and ICASSP.
- 99% success rate on zero-shot laundry: Sunday Robotics debuts their new ACT-2 model (36 points · r/singularity · discussion) -- Sunday Robotics has debuted its ACT-2 model, which achieves a 99% success rate on zero-shot laundry folding tasks.
- Lawsuit Claims the Mayo Clinic's Use of AI Is Butchering Patient Care (33 points · r/ArtificialIntelligence · discussion) -- A lawsuit claims the Mayo Clinic's use of AI is harming patient care, raising concerns about AI deployment in healthcare settings.
- User experience of Bonsai-Ternary-27B on 4060Ti 16GB for KB management and productivity assistant use cases (32 points · r/LocalLLaMA · discussion) -- A user shares their experience running Bonsai-Ternary-27B on a 4060Ti 16GB for knowledge base management and productivity tasks, finding it handles complex setups with Obsidian vaults, project management integrations, and memory systems reasonably well as a local alternative to cloud models.
- Why is ECCV so insanely expensive for students presenting papers? (31 points · r/MachineLearning · discussion) -- A student researcher expresses shock at ECCV registration fees—$440 for early bird students and $805 for full registration—and notes that presenting a paper requires a full registration, effectively punishing students for having their work accepted. Travel grants and registration waivers have been rejected.
- A new, state-of-the-art, agentic pipeline: concept + track = full audiovisual world (30 points · r/artificial · discussion) -- A new agentic pipeline has been shared that combines concept and track generation to produce full audiovisual worlds, representing a significant advancement in AI-generated content creation.
- Will we get accessible open-source models again? (28 points · r/LocalLLaMA · discussion) -- A user expresses concern that all recent open-source LLMs are in the 0.5T+ parameter range, questioning whether Chinese labs will ever release smaller 20B-120B models again given their focus on benchmark competition and low API pricing rather than serving individual users.
- Anthropic Leads top 10 models by $/spent (27 points · r/OpenAI · discussion) -- A chart showing Anthropic's models leading the top 10 in cost per token generated discussion about the economics of frontier AI models and whether the premium for Anthropic's models is justified by their quality.
- Prism accidentally leaked (26 points · r/MachineLearning · discussion) -- Prism, a free alternative to Overleaf that was formerly called Crixet and is now owned by OpenAI, accidentally leaked another user's paper.
- AI Wealth-Sharing Plans Gain Support From Trump and Sanders (17 points · r/ArtificialInteligence · discussion) -- Wealth-sharing proposals related to AI-driven job displacement have gained unexpected bipartisan support from both Donald Trump and Bernie Sanders, reflecting growing political recognition of AI's economic impact.
- Gemini is EVERYWHERE (10 points · r/artificial · discussion) -- Gemini is integrated into Chrome, Google Search as AI Mode, Android phones, Circle-to-Search, Google Maps, Gmail, and Docs, creating an ecosystem compatibility comparable to iOS.
- Gpt-5.6 and Grok 4.5 dropped in the same 24 hours and my Slack was quieter than i expected (9 points · r/ArtificialInteligence · discussion) -- A developer notes that the simultaneous release of GPT-5.6 and Grok 4.5 generated surprisingly little workplace excitement, suggesting that the rate of meaningful model improvement has slowed and teams are settling into incremental tuning rather than transformative upgrades.
- Genie 3 Isn't About Soulless Games, It's About Whether Creative Craft Careers Survive the 'Vibes' Metric (6 points · r/artificial · discussion) -- An analysis of Google Genie 3's ability to generate explorable game worlds from text prompts examines the broader implications for creative craft careers like level design and environmental art, questioning whether the technology will expand small team capabilities or drive massive headcount cuts at studios.
- Finally, an AI start up with a Billion-dollar revenue not valuation (backed by Nvidia) (6 points · r/artificial · discussion) -- An AI startup backed by Nvidia has reportedly reached $1 billion in actual revenue rather than just valuation, marking a rare milestone of profitability in the AI industry.
- AI Executives Add Personal Security as Backlash Turns Violent (4 points · r/ArtificialIntelligence · discussion) -- AI executives are drastically increasing personal security after a 20-year-old attempted to murder Sam Altman with a Molotov cocktail and gunfire, while organized anti-data center opposition groups doubled to 833 across 49 states and blocked 75 projects worth $130 billion in Q1 2026.
- Wan Team's WanSong Generates 5-Minute Songs With Dual Stems (4 points · r/ArtificialIntelligence · discussion) -- The Wan Team released WanSong, a pure diffusion-based music model that generates full songs up to five minutes long in a single pass with separate vocal and instrumental stems, using step-distillation for faster inference and supporting fine-tuning for downstream editing tasks.
- My AI agents have now run on four model generations. Their memory never noticed. (2 points · r/artificial · discussion) -- A user reports running the same persistent AI agents across six different model generations (Sonnet 4.5 through Claude 5 family), with the agents maintaining continuity through JSON and markdown on disk rather than model-specific capabilities, treating model swaps as a simple config line change.
- Anthropic IPO Could Launch in October as China's Kimi K3 Overtakes Claude (2 points · r/artificial · discussion) -- Reports suggest Anthropic's IPO could launch in October, potentially timed around Kimi K3's overtaking of Claude on performance benchmarks.
- Moonshot's Kimi K3 sends AI and semiconductor stocks into a tailspin (1 points · r/artificial · discussion) -- Moonshot's Kimi K3 has sent AI and semiconductor stocks into a tailspin, reviving DeepSeek-era fears about the economics of US infrastructure spending.
- Requential Coding. Researchers achieved <1 bit compression due to the generalization ability fostered by advanced teaching technique (1 points · r/artificial · discussion) -- Researchers achieved less than 1 bit compression in coding through a generalization ability fostered by an advanced teaching technique.
- Compiled three main AI incident databases into one readable digest. (1 points · r/artificial · discussion) -- A user has compiled three main AI incident databases into a single readable digest for easier tracking of AI-related incidents.
- Chiron: Exact recovery + held-out verification + refusal engine (public repo + prototype live) (1 points · r/artificial · discussion) -- Chiron is a verification system that recovers exact underlying rules via Minimum Description Length, proves them on held-out data, and refuses to stamp anything it cannot verify exactly, with signed falsifiable certificates.
- I made something that feels like GPT-Live but you can run it yourself in Rust (1 points · r/artificial · discussion) -- A developer has built a GPT-Live-like voice assistant that can be self-hosted, implemented in Rust.
- Agentwashing Sounds Like a SaaS Product. It's Actually a Confession. (0 points · r/artificial · discussion) -- VentureBeat's Pulse Research found that 71% of enterprises admitted a quarter or fewer of what they called agents were real multi-step workflows, with only 10% having more than half their fleet crossing that bar, while 27% had no real-time visibility into agent runtime costs.
- the hard part of a private ai rollout isn't the vpc, it's the connectors nobody wrote (0 points · r/artificial · discussion) -- The real challenge of private AI rollouts is not the VPC deployment but the downstream connectors for internal systems like billing tools and ops dashboards, with deployment taking about a week while mapping internal systems into something an agent can call is the actual project.
- If AI disappeared tomorrow, what part of your workflow would be affected the most? (0 points · r/artificial · discussion) -- A discussion thread asking what part of users' workflows would be most affected if AI suddenly disappeared, with common answers including debugging, summarizing documentation, brainstorming, and writing SQL and boilerplate code.
- 1 Person + AI + Email Automation = A Successful Web Agency (0 points · r/artificial · discussion) -- A solo developer describes running a full web agency using AI to build websites and email automation tools like Swokei to find leads and send personalized outreach about website redesigns.
- What Is Left for Us to Become? (0 points · r/artificial · discussion) -- A philosophical essay examining whether humans relate to AI as instruments that extend reach or as oracles that receive authority, arguing the danger lies not in machine capability but in the posture humans bring to the exchange.
- Meta reins in new AI tool after criticism (0 points · r/artificial · discussion) -- Meta has scaled back a new AI tool following public criticism.
- AI Is this legal? (0 points · r/artificial · discussion) -- A user questions the legality of an email from a financial advisor that included an AI-generated response suggesting follow-up questions.
Updates: 05:30 AM PDT · 07:24 AM PDT · 08:30 AM PDT · 11:30 AM PDT · 02:30 PM PDT · 05:30 PM PDT