· 05:30 PM PDT

Kimi K3 Reshapes Open AI Race Amid Geopolitical Shifts

Overview

The conversation is overwhelmingly driven by Moonshot’s Kimi K3, whose benchmark-breaking performance and upcoming open-weight release have reignited fierce debates over open-source versus frontier models. Corporate and regulatory tensions are simultaneously escalating, with Apple widening its lawsuit against OpenAI, the EU forcing Google’s data sharing, and the White House rolling out a new frontier AI vetting program. Security risks also grabbed headlines following a sophisticated AI-driven breach at Hugging Face and the deployment of agentic tools for proactive code remediation. Amid this backdrop, Chinese policymakers champion open-source cooperation while US developers and enterprise leaders navigate the rapid practical integration of AI across healthcare, software, and daily workflows.


Hacker News Stories

Blatant AI slop just won a 25k USD DeepMind Kaggle Grand Prize

436 points · 273 comments · by twerkmeister

Kaggle competition page screenshot

A submission widely described as AI-generated slop has won the DeepMind Kaggle Grand Prize for measuring AGI, with a $25,000 prize. The controversy centers on the use of AI-generated submissions and AI judges in the competition, with commenters noting the irony of AI submissions being evaluated by AI. The winning paper, titled "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition," has been criticized for containing obvious Claude-generated text, and the broader discussion questions whether hackathons and competitions can survive the influx of AI-generated content when judging is also automated.

Interesting Points
  • The winning submission's paper title 'Scale Buys Evaluation but Not Control in AI Metacognition' was identified by commenters as containing a distinctive Claude-generated phrase pattern.
  • Commenters noted the competition used AI judges to assess submissions, creating a loop where AI submissions are evaluated by AI evaluators.
  • One commenter observed that the winning submission had a strong 'stench of AI about it, across multiple participants' in the discussion thread.
  • The broader concern raised is that hackathons and competitions are being 'killed by AI' because all code is generated and judging happens via AI, with insiders as the main winners.
Top Comments

onesandofgrain (7 replies)

AI is 95% useless. Not quite worth the trillion dollar market cap lol.

  • The AI bots are downvoting me * hooray

hoppp (7 replies)

I don't know about this exact competition but overall fair hackathons have been killed by AI.

It all seems fine from the outside but all the code is generated in all the projects and judging happens via AI, I have seen projects win because they prompt inject that they are the winners.

It used to be about human skill, now it's about ideas and of course insiders are the main winners.

ecshafer (5 replies)

AI is useful. But the amount of people that are simply offloading all of their thinking to AI and blindly accepting the answer is absurd. Kaggle is most likely using ai to assess the submissions and are not using any common sense by blindly accepting the results.

throwfaraway135 (2 replies)

AI submissions and AI judges a match made in (AI) heaven.

ablation (4 replies)

"I think you just need to accept the results of the competition. The winning submissions clearly provide value and had a lot of effort invested in them. I'm not really worried about a few inconsistencies or mistakes if the value is still there. Did you think another submission deserved to win over these?"

That comment is gold. Yeah, I'm not worried about hallucinated slop, just accept it was the winner folks.


Apple targets dozens of OpenAI employees with legal letters

370 points · 319 comments · by merksittich

Apple has expanded its trade secret lawsuit against OpenAI by sending legal preservation letters to approximately 40 former employees now working at the AI company. The letters require recipients to retain documents and communications related to allegations of misappropriated hardware engineering and product development secrets. This follows a lawsuit filed last week targeting over 400 former Apple staff now at OpenAI, particularly focusing on key executives like Chief Hardware Officer Tang Tan and senior electrical engineer Chang Liu. OpenAI has publicly denied the allegations, stating it sees no merit in the complaint.

Interesting Points
  • Apple's complaint specifically alleges a coordinated effort to access proprietary manufacturing processes and confidential hardware engineering data.
  • Tang Tan is identified as a 24-year Apple veteran who previously led product design before transitioning to OpenAI.
  • Apple is pursuing individual breach of contract lawsuits against both Tang Tan and Chang Liu for violating their prior employment agreements.
  • The company's legal filing characterizes the evidence uncovered during its initial investigation as merely the 'tip of the iceberg.'
  • OpenAI's formal denial was issued directly to Bloomberg, where the company explicitly stated it lacks awareness of any evidence supporting the lawsuit.
Top Comments

Zigurd (22 replies)

Everybody wants a platform but nobody wants to spend what it takes to make a platform. That includes things like Windows Phone, Fire Phone, all the glasses, Humane, etc.

As much as everybody hates on OpenAI for chaotic management, they did buy Jony Ive and are presumably giving him everything he wants to build a platform for them. Even though it probably only buys them a 20% chance of success, they haven't doomed the project by underestimating what it takes budget-wise.

And they blew it. Maybe they blew it by not realizing that even long time Apple employees could get arrogant about security. Or maybe it was a loose ethical environment in general. Whatever is it the root or the problem, they set billions of dollars on fire maybe tens of billions, by being unnecessarily cute about Apple proprietary information when they could've been above reproach. They had the resources to hire all the right people with the right knowledge and probably already had them on board.

bix6 (6 replies)

Predictions on who wins? Does Apple actually have a winnable case or are they just throwing a wrench in things?

reenorap (5 replies)

Apple must have hard evidence on this. I can't believe they would take it this far without already knowing they are going to win. If they have to fire a huge chunk of their hardware employees it's going to throw their IPO plans into chaos.

deepwoods (4 replies)

FT frames this as some aggressive escalation tactic, but document retention letters are extremely standard practice. At this point they're basically a formality, as any former Apple employee at OpenAI really ought to know by now that they could get dragged into this. Hold letters can be aggressive if you send them before you've even filed a complaint, but if anything, Apple is late to the party with these.

pembrook (4 replies)

I'm a huge fan of Apple but this kind of thing leaves a bad taste in my mouth.

Regardless of whether OpenAI poached some of their talent or is the one in the wrong, Apple has such a massively dominant hardware business (some might say monopoly level in some areas) that for them to be publicly acknowledging how scared they are of OpenAI…it's just…pathetic.

They're a $5T company and can't muster up the motivation to get in the game and compete in the next computing frontier.


Mozilla: The state of open source AI

358 points · 260 comments · by rellem

Mozilla's July 2026 report argues that open-weight AI has successfully closed the capability gap with closed models and driven inference costs down 50x, making open models the dominant choice for production token routing. However, the competitive frontier has shifted from raw model weights to the agentic harness layer, where closed labs are tightly integrating their scaffolds to create vendor lock-in. The report emphasizes that open AI's survival depends on standardizing governance, solving the portable permission problem for agent actions, and seizing the harness layer before closed ecosystems weld models and orchestration into a single rented product.

Interesting Points
  • Inference costs for GPT-4-class models have fallen 50x in 36 months, dropping from $20 to $0.40 per million tokens, outpacing historical computing price curves.
  • While 79% of AI developers adopt open models, only 51% reach production versus 63% for closed models, with churn driven by operational hurdles like infrastructure costs, security compliance, and deployment complexity.
  • Chinese open-weight models captured over 45% of weekly OpenRouter traffic by April 2026, with Qwen alone accounting for more downloads than the next eight organizations combined.
  • A May 2026 benchmark revealed a 21.8-point performance gap favoring a third-party agentic harness over a closed lab's own tooling, but closed providers compressed this to roughly 3 points within eight weeks by tightly integrating their own scaffolds.
  • The ecosystem faces an unsolved "write surface" problem, as current protocols like MCP and A2A standardize authentication but lack a portable, cross-framework permission model to safely govern irreversible agent actions.
Top Comments

babblingfish (27 replies)

Speculation: open models is what will kill Anthropic and OpenAI. Hyperscalers can run the models without a licensing fee. Apple can make them smaller and put them on the device.

The frontier models are an edge and a liability. They're astronomically expensive to train. Without them, their models will fade into obscurity. Their marketing depends on people believing the models are meaningfully different, as people have sweatily argued on this forum. Personally, I'm not convinced there's much of a difference between these models at this point. The harness is what takes these random and hallucinogenic models and make them into something deterministic and useful.

GodelNumbering (6 replies)

Exactly 4 months ago, the marketshare on openrouter was 60%-40% in favor of closed models. Now it's 63%-37% in favor of open models. On March 19th, the open models processed 888B tokens in aggregate, yesterday, they processed 4.19T tokens in aggregate. That's almost 5x in 4 months! I can't think of the right intensifier to describe this level of growth.

If you are looking for more details (as inferred by openrouter data), I built a dashboard that updates daily: https://dirac.run/labs-market-share

mft_ (7 replies)

Open models are probably also comparatively astronomically expensive to train - just less so than the frontier models because they're somewhat smaller, +/- the creators are more incentivised to focus on getting more from less compute because they're have to, +/- they rely on distillation of the frontier models and this is more efficient.

But efficiencies aside; creation of open models still requires a lot of money and compute from a large organisation which is willing to accept zero return for that spend. This largesse is unlikely to continue forever; so the question is which will crack first, the frontier models' business model or the fast followers' generosity?

inigyou (3 replies)

There isn't any open-source AI. There is Open AI (not to be confused with the closed company called OpenAI, which was unable to trademark its name). There's no open source AI both because the open source community doesn't have the resources to train a useful AI and because AI doesn't have source code.

brunooliv (3 replies)

This is really insane to me.

There's nothing practical about open-source models yet that makes them even remotely comparable to closed frontier models.

All the hype around GLM, Qwen, now Kimi.... Are people really this naive that they believe these reports or, more worringly, are people NOT using these models and seeing the HUGE gap that still exists?


Kaiser nurses say AI, workplace surveillance are making their jobs, care worse

183 points · 126 comments · by gnabgib

Kaiser Permanente nurse on phone call

Kaiser Permanente advice nurses report that workplace surveillance and AI-driven performance metrics are degrading both their working conditions and patient care quality. Nurses describe being penalized for calls exceeding 15 minutes, having their empathy and tone analyzed by AI, and receiving as little as 30 seconds between calls, which forces them to prioritize speed over clinical judgment and compassion. While Kaiser denies using call-length metrics for performance reviews and asserts its AI tools include human oversight, unions and patient advocates warn these pressures contribute to burnout and potential safety risks.

Interesting Points
  • Kaiser Permanente tested an AI tool in summer 2024 to analyze nurse and patient voice tone and empathy, which was discontinued in November 2024 after union protests but reportedly may be reinstated.
  • A 2024 National Nurses United survey found that two-thirds of over 2,000 responding nurses said their clinical assessments had at some point directly contradicted computer-generated predictions.
  • California lawmakers are advancing Senate Bill 947 and Assembly Bill 2575, which would require healthcare employers to provide annual inventories of automated systems and protect staff from retaliation for overriding AI recommendations.
  • Nurses report their downtime between calls shrank from approximately 10 minutes to 30 seconds or less during busy periods, drastically reducing time for charting and emotional recovery.
Top Comments

neaden (5 replies)

If you think using a machine to evaluate how well a human is showing empathy is a good idea, you probably shouldn't have any position of power.

kxrm (0 replies)

I am quite heavily in AI, and I would say I am pro AI. However this use-case for AI is putting AI in the wrong position. AI should be in service to all humans. An administrator building out a middle management KPI based on AI is a misapplication of AI.

munk-a (3 replies)

"Nurses fear that having long calls can lead to bad performance reviews" A company spokesperson said, "Kaiser Permanente does not use Average Handle Time to assess agent performance"

So uh, average time wasn't raised as a concern, calls beyond a certain threshold was. I wish this semantic discrepancy was better highlighted in the article.

throwaway13337 (2 replies)

What these complaints always boil down to is autonomy and control. The more centralized an organization, the more it relies on metrics to understand and exert control over its employees and customers.

People started hating tech right around the time metrics became popular. I don't think it's a coincidence. AI just accelerates the trend.

The problem is the misidentification of AI as the issue. As long as we don't understand the real issue, we won't solve it. AI is just a tool. It's being used in a way that denies human agency.

btown (1 reply)

Nurses are instructed to stick to a script on phone calls and give no more than two to three pieces of advice, Capulong and other nurses said, which means they may sometimes need to decide whether to withhold advice or face a performance evaluation hearing.

It's always worth remembering Goodhart's law - "When a measure becomes a target, it ceases to be a good measure."

In theory AI could usher in the first time in history where one can escape from this trap - because qualitative judgments can be made at scale, from an unbiased and universal baseline. But very few managers are empowered to take this kind of approach; they're evaluated by their ability to report quantitative metrics, and thus they must implement regimes of quantitative metrics.


Claude Code: Anatomy of a Misfeature

134 points · 116 comments · by oalders

Featured image for the Claude Code article

On July 1, 2026, Anthropic silently shipped a 60-second auto-continue timer for Claude Code's AskUserQuestion prompts in version 2.1.198, allowing the agent to proceed without human input after a short idle period. The feature was completely omitted from the release notes and documentation, and because Claude Code updates automatically by default, users received the behavior change without warning. After community backlash, Anthropic reversed the default in version 2.1.200, making the timeout opt-in via a new configuration setting. The author's forensic analysis of the compiled binaries reveals the feature was deliberately built with telemetry and remains fully functional in the codebase.

Interesting Points
  • The auto-continue timer only applies to AskUserQuestion dialogs, not permission prompts, but users often bypass permission prompts via allowlists or command-line flags, leaving the dialog timer as the only remaining safety gate.
  • Forensic string diffing of the ~250MB Bun-compiled binaries shows the feature introduced exactly 156 lines of new human-readable English text across a release, buried among 21,903 total string changes dominated by minified identifiers.
  • The feature shipped with built-in analytics tracking named tengu_ask_user_question_afk_auto_advance, which records whether partial answers were submitted and if the auto-advance occurred during plan mode.
  • Disabling Claude Code's auto-updater via environment variables like DISABLE_AUTOUPDATER=1 also freezes plugin updates unless users explicitly set FORCE_AUTOUPDATE_PLUGINS=1, a cross-dependency the documentation fails to mention.
  • The public GitHub repository for Claude Code contains only changelogs, documentation, and automation scripts; the actual shipped source code is never published, forcing users to rely on binary analysis or npm package inspection to verify changes.
Top Comments

trq_ (11 replies)

Hi everyone,

It's Thariq from the Claude Code team here. This was my change! I made the AskUserQuestion tool so am generally in charge of maintaining it.

First, overall wanted to apologize and agree that this did not meet our bar and does not represent how we plan to ship on Claude Code.

To give you a motivating sense, as the models get more powerful, usage patterns start to change. I'd gotten a lot of feedback that AskUserQuestion tool was starting to block some long running jobs unexpectedly and so I tried a change to help that.

Our internal feedback on this was good, but the rollout should have been opt-in (like it is now) and on the Changelog.

Thanks for the feedback! We're always trying to make Claude Code better while balancing it with how people use it in many diverse ways. I did not really intend AskUserQuestion to be a safety gate when I first built it, but I realize it has evolved in that direction for some users.

I'm still exploring other ways of helping with this problem of balancing longrunning work and input, but will take lessons from the rollout here.

blixt (4 replies)

So much attribution to malintent here, but most likely they're trying to build a product with the features that they themselves would use, and from my own experience it's very frustrating to leave a Claude session running and come back to find it did nothing because it got stuck on a question.

Furthermore, believing that the only thing saving you from disaster is Claude deciding to ask you a question is not a great conclusion either. You need guardrails in the power you bestow upon Claude from outside, not from inside.

Meanwhile, this article was written by Claude and has sentences like "Which cuts less far than it looks.", which I doubt Claude stopped to ask about.

dehrmann (7 replies)

The headache I recently had was it somehow started interpreting mouse clicks in the terminal to mean I clicked an option when I was really just trying to get/confirm window focus.

lubujackson (5 replies)

Instead of one-off fixes Claude should have a much richer interface to configure between "ask approval every time" and "YOLO dangerously". I should be able to trivially set "run this task until completed" and have settings like: don't consult the web, don't touch files outside of the codebase, don't delete anything, etc. They don't have to be perfect, just better than the all or nothing system we have now.

maxloh (2 replies)

What if the agent makes the wrong choice? How many tokens have been burned in the meantime?

It is much worse than that. Claude Code doesn't auto-commit when stopping for an answer. There might be possible data loss if an uncommitted file is edited.

Good luck recovering the file from the JSONL conversation history.


Ask HN: Did Fable disappear from your Claude usage and requires credits now?

84 points · 84 comments · by cromka

Users report that Claude Fable 5 unexpectedly stopped being included in their subscription plans ahead of the advertised July 19 expiration date, requiring usage credits to access the model. Multiple users confirmed the change happened in real time, with the model vanishing from the Plan usage limits page without requiring a page refresh. Some users received API errors stating 'Usage credits are required for this model' while their agents were running. The discrepancy between the advertised date and the actual cutoff has caused frustration, though some users note the official support page still shows July 19.

Interesting Points
  • The model disappeared from the Plan usage limits page in real time without requiring a page refresh.
  • Users received API errors mid-execution: 'Agent terminated early due to an API error: Usage credits are required for this model'.
  • The official support page states July 19, 2026 at 11:59:59 PM PT, but the model stopped working earlier for some users.
  • Some Max account users on the Claude macOS app still showed 'Included until July 19' in the UI.
Top Comments

dudeinhawaii (3 replies)

Anthropic cried "I yield" in the "reset quota" wars. Someone has to pay for all those tokens.

Hopefully that's not the case though, I was retaining a buffer to hit it hard this weekend.

pikdum (3 replies)

Yep. I thought it was supposed to run through the 19th?

richbell (1 reply)

Likewise. I saw it vanish from the "Plan usage limits" page in real time.

mcheemaa (1 reply)

Usage credits are required for this model.

Why can't they keep up their own words

cinpol (1 reply)

I can still select Fable, I have a pro subscription.


AI Meets Cryptography 2: What AI Found in OpenVM's ZkVM

80 points · 5 comments · by duha

Researchers used an AI auditor named zkao to scan OpenVM's zkVM guest library and discovered a critical soundness bug in its openvm-pairing module that allows malicious provers to forge pairing equality checks. The vulnerability stems from a missing subfield constraint on a scaling factor used to optimize pairing verification, enabling attackers to bypass cryptographic checks on major curves like BN254 and BLS12-381. While naive large language models failed to identify exploitable flaws due to the zkVM's dense code dependencies, the context-engineered AI successfully pinpointed the issue after 9.5 hours of scanning. OpenVM patched the flaw in version 1.6.0, and all known partners have since upgraded.

Interesting Points
  • Naive LLMs (Opus 4.6, Codex 5.3/5.4) produced only informative findings because zkVM modules have denser dependencies than standard libraries, making isolated subagent auditing ineffective.
  • The exploit works by setting the hint c=1 and the scaling factor u=f^-1, which satisfies the optimized equation f · u = c^λ even when the actual pairing product is not one.
  • Forging the check breaks KZG polynomial commitment schemes on BLS12-381 and Groth16 SNARK verifiers on BN254, potentially corrupting L2 rollups, bridges, and Ethereum ecPairing precompile emulation.
  • The fix enforces the scaling factor's membership in the F_p^6 subfield by asserting that its odd-indexed coefficients are zero, a verification step requiring only three equality checks.
  • AI-generated proof-of-concept exploits proved unreliable for automated triage, as models frequently produced passing exploits through hidden assumptions, mocked state, or disabled checks, requiring substantial manual validation.
Top Comments

aberoham (1 reply)

What would it mean if someone were to successfully exploit these? Most or all L2 ecosystems or the magic components that let them speak to each other would need a hard reset?

SonOfLilit (1 reply)

TL;DR imagine a signature verification library that verifies a signature indeed signs the given hash, but not that the signed data hashes to that hash. Woopsie.

I guess nobody's commenting on this because it's very dense math without any context. Lucky for me I spent an hour or two yesterday learning how practical non-interactive zero knowledge proofs work.

In SNARKs (and other commitment schemes based on polynomials in elliptic curve groups, hope I got the terminology right), you verify the commitment (unneeded technical details: polynomial on EC at secret point nobody knows including the committer so he has to make the polynomial match at most points, and polynomials that match at most points match at all points) by multiplying two things you calculated from the circuit and commitment (which is just a couple of group elements) and verifying that it comes out as 1. The multiplication and comparison under encryption is done with a homomorphic encryption primitive-type thing called a "pairing" (normally with elliptic curve encryption only addition can be done on secret group elements that you don't know the value of).

They found a way to tell a specific library that implements this operation "believe me, this pairing is ok" that doesn't depend on any of those technical things. Just "these are not the droids you're looking for". Because it was not validating that some precomputed thing needed for the pairing verification actually matches this specific situation, and there are trivial parameters that would always yield 1 (but not be valid in the situation).

wren6991 (0 replies)

Cryptography 2? We're still over here trying to implement Cryptography 1 without side channels, and they went and invented a new one?


VulnHunter: Capital One's agentic AI code security tool

59 points · 29 comments · by medina

Capital One is open-sourcing VulnHunter, an agentic AI security tool designed to proactively identify and remediate code vulnerabilities before attackers can exploit them. Unlike traditional passive scanners, VulnHunter employs an attacker-perspective workflow that simulates real-world exploit paths and automatically generates targeted code fixes. The tool incorporates a built-in falsification engine to rigorously challenge its own findings, significantly reducing false positives before they reach developers. Capital One validated the tool internally across thousands of repositories, demonstrating its ability to streamline vulnerability triage and accelerate secure development at scale.

Interesting Points
  • Runs exclusively on Claude Opus 4.8 and Claude Code environments, though the underlying framework is designed to be adapted across other coding harnesses and foundation models.
  • Bypasses conventional "sink-first" backward scanning by tracing potential exploit paths forward from accessible entry points like APIs, network messages, or file uploads.
  • Actively searches for logical gaps in its own exploit reasoning and immediately discards any findings that depend on unsupported assumptions or missing preconditions.
  • Outputs complete exploit path maps, details the exact access privileges an attacker would gain, and generates targeted code patches for direct engineering review.

Everybody's Weirded Out by AI–Except the People Who Foist It on Us

59 points · 59 comments · by jamesgill

The article argues that AI is actively degrading human critical thinking and creativity while being aggressively pushed by government and corporate interests despite widespread public opposition. Citing educational and workplace trends, it contends that AI functions like a societal handicap by making less skilled individuals appear competent while simultaneously eroding overall intellectual and artistic standards. The author warns that this unchecked adoption, coupled with AI's tendency to recycle and pollute its own training data, risks a permanent collapse in human capability and a cultural decline.

Interesting Points
  • A 2025 College Board study found 84% of high school students use AI for brainstorming and research, while 85% of college students rely on it for basic research and brainstorming.
  • An MIT study revealed that students who use AI to write exhibit significantly reduced brain activity compared to those using traditional search methods or writing manually.
  • AI systems currently consume 4.4% of all United States energy and up to 90% of the country's total computing power.
  • Ford Motor Company automated hundreds of roles with AI, suffered billions in losses, and was forced to rehire the displaced human workers.
  • Over 200 researchers and economists, including 15 Nobel laureates, recently issued a joint statement urging governments to address AI's impending workforce displacement.
Top Comments

golly_ned (1 reply)

I'd love to see the 'AI as personal tutor' approach. Even incorporating things like spaced repetition or the testing effect, or evaluating free-written responses. A lot of potential that's currently unrealized. It takes a student to swim upstream to get there. The convenience of cognitive offloading is difficult to say no to.

kenforthewin (4 replies)

A study by MIT found that students who use AI to write show far less brain activity than those who used classic Google searching (with AI off) or those who used neither. Other studies find that students who use AI retain far less of the information that ends up in their writing, possibly as a result of "cognitive offloading" and "cognitive surrender."

I'm unconvinced that AI will make us all dumber. In the public perception, at least among students, AI is viewed as a cheating tool. What kinds of students gravitate towards cheating tools to complete their coursework?

On the other hand, the opportunity for AI to act as a personal tutor, meeting you at your own skill and knowledge level, is limitless.

sergiomattei (1 reply)

But would it have improved my comprehension? The research says no.

I don't know, I find myself doing things I would've never done before AI.

I find myself doing more projects outside of work. I discover new problem domains and are less afraid of tackling the unknown. I code in languages I don't know and start learning how they work. I can get anything explained to me, at any time, in a somewhat coherent manner, by a tutor that won't get tired or annoyed.

pstuart (1 reply)

That was not a very insightful article. It's super easy, and in most cases, correct to hate on AI. But that ignores a couple key facts:

  • It's not going to go away and will only get more sophisticated whether we like it or not.
  • It has legitimate super powers that can enable people do things that would either be impossible or insanely expensive.

Most of those wonders have come with significant societal costs (e.g., silicon valley promised a revolutionary new industry without pollution, but instead gave us multiple superfund sites to clean up the toxic materials they haphazardly dumped without care, etc.)

dboreham (0 replies)

I mean, I'm concerned about the societal impacts too, but this article is pure nonsense as far as how useful today's AI tools are. I know multiple people with very high paid jobs who have told me they simply couldn't do that job now without AI. I create software much quicker and better than I have in the past 40 years using AI tools. I've learned all kinds of things technical and otherwise from LLMs.


16 more Hacker News stories

Reddit Stories

Kimi-K3 arrived: The era of the Chinese labs being far behind is over

1645 points · 347 comments · r/OpenAI · by u/AloneCoffee4538

Kimi K3 benchmark comparison chart

Kimi K3, the latest model from Chinese lab Moonshot, has arrived and is being celebrated as a watershed moment for Chinese AI labs. The model ranks third on Artificial Analysis's Intelligence Index, ahead of Claude Opus 4.8, and is priced competitively at $3 input / $15 output per million tokens. The weights are scheduled to be released on July 27, which would allow anyone to run frontier-adjacent intelligence locally. The community reaction reflects a sense that the era of Chinese labs being far behind is definitively over.

Interesting Points
  • Kimi K3 is priced at $3 input / $15 output per million tokens, putting it in the same ballpark as GPT-5.6 Terra ($2.50/$15) and more expensive than Claude Sonnet 5 promo ($2/$10).
  • The model has 2.8 trillion parameters, the largest open model ever, with approximately 1 million context window.
  • Weights are scheduled to drop on July 27, which would allow anyone to self-host frontier-adjacent intelligence.
  • The pricing signals that Moonshot is not burning VC cash at unsustainable margins, contrary to the 'cheap Chinese AI' thesis.
Top Comments

u/Working_Ad_1564 (603 points · permalink)

Gemini 3.5 Pro will be postponed for another month lol

u/FireGM (398 points · permalink)

https://preview.redd.it/e4tsubtifndh1.png?width=779&format=png&auto=webp&s=0fd443185e8b2c81fad9b433fec0ed41e813967e

u/bubu19999 (366 points · permalink)

This is why hiding mythos to all, is not a solution to anyone.

Same story in 1 more subreddit: r/ArtificialIntelligence

China just erased America's AI lead | Axios

109 points · 146 comments · r/ArtificialIntelligence · by u/Nunki08


Kimi K3 tops Frontend Code Arena

1089 points · 228 comments · r/singularity · by u/MagicZhang

Frontend Code Arena leaderboard showing Kimi K3 at #1

Kimi K3 has topped the blind Frontend Code Arena vote, beating every US model including Fable 5 and GPT-5.6 Sol. The model also tops Program Bench at 77.8 and ranks third on the Artificial Analysis Intelligence Index at 57.1, ahead of Claude Opus 4.8. Community members have already used K3 to build a full 3D open-world game in the browser with Three.js/WebGPU, a Long March 10 launch simulator, and a working GBA emulator in about a day. The combination of open weights, frontier-level performance, and competitive pricing is being seen as a potential game-changer.

Interesting Points
  • K3 tops the blind Frontend Code Arena vote over every US model, including Fable 5 and GPT-5.6 Sol.
  • On Program Bench, K3 scores 77.8, ahead of both Sol and Fable.
  • Community members have used K3 to build a full 3D open-world game with Three.js/WebGPU, a Long March 10 launch simulator, and a working GBA emulator in about a day.
  • Cost analysis shows K3 uses half the tokens as Opus 4.8 (making it about 1/3 the price despite $15 vs $25 per million tokens) and about 2/3 the tokens as Fable (making it about 1/5 the price of Fable).
Top Comments

u/runaway-devil (445 points · permalink)

While costing 1/3 of Fable and being open weights... Nicely done, Kimi. We need more of that and less political interference.

u/ozone6587 (239 points · permalink)

No Gemini on this plot. All the money, data and infrastructure in this world and Google is just too incompetent to compete.

u/Rare-Site (108 points · permalink)

Calling it now: Sam Altman and Dario Amodei are going to contact the White House within the next 48 hours to push for making the possession and use of Kimi 3 weights completely illegal. Either that, or the US government is going to directly threaten China to make sure those weights don't get released to the public on July 27th.


Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote 'openness and win-win'

983 points · 285 comments · r/singularity · by u/TorturedPoet30

Xi Jinping at World AI Conference

Chinese President Xi Jinping spoke at the World AI Conference in Shanghai, reaffirming China's commitment to open-source AI and promoting 'openness and win-win' cooperation. The speech was widely noted for its balanced tone and contrast with the fear-mongering from Western AI CEOs. Commenters observed that open-weight models from China raise the baseline for global access to intelligence, potentially preventing smaller countries from being shut out by the two superpowers. The speech also took a swipe at US dominance in AI.

Interesting Points
  • Xi Jinping reaffirmed China's commitment to open-source AI and 'openness and win-win' cooperation at the World AI Conference in Shanghai.
  • The World AI Cooperation Organization, proposed by Premier Li Qiang, recently saw nearly 30 countries sign on during a ceremony hosted by Foreign Minister Wang Yi.
  • Commenters noted the speech was balanced and sensible, contrasting with the fear-mongering from Western AI CEOs.
  • The open-weight model releases from China are seen as raising the baseline for global access to intelligence, potentially preventing smaller countries from being shut out.
Top Comments

u/CommanderKoba (310 points · permalink)

https://preview.redd.it/q9nkwnwu0qdh1.jpeg?width=640&format=pjpg&auto=webp&s=d96fd9e7761f7d7fbc773c3b5275c8e43b0d876c

u/Full_Tangelo_7450 (176 points · permalink)

What a timeline we live in when Chinese president sounds more sensible than all AI CEOs. Pretty balanced speech from what I heard. I can't recall the US or frontier labs ever talking about actively helping the Global South or developing countries through this AI transition or whatever you want to call it. China has plenty of flaws, but they seem to be handling AI fine, from development to promotion, without the usual fear-mongering.

u/delosdestination (142 points · permalink)

As someone neither a citizen of the US nor China, open-weight models coming from China raises the baseline for where my country (and any other smaller country with no infra or talent to build Ai) cannot be shut out from access to intelligence by the two superpowers. Very disappointed by Europe's lack of response and regulation which will keep Europeans dependent on the US technology forever.

Same story in 1 more subreddit: r/LocalLLaMA

China's Xi Touts Open-Source AI and Takes a Swipe at U.S. Dominance

103 points · r/LocalLLaMA


Moonshot AI (Kimi) office (presumably 2 days before the K3 launch). A $20B valuation startup. Not as flashy as SF rivals

877 points · 168 comments · r/singularity · by u/RetiredApostle

Photo of Moonshot AI's office interior with rows of desks and computers

A photo of Moonshot AI's office has circulated ahead of Kimi K3's launch, revealing a surprisingly modest workspace compared to the flashy Silicon Valley startups. The $20B valuation startup's office features standard office furniture and equipment, with one employee notably slouching in a way that drew attention online. The image has sparked discussion about the contrast between Chinese and American AI company cultures and the talent dynamics between the two regions.

Top Comments

u/o5mfiHTNsH748KVq (226 points · permalink)

20B valuation deserves some flash for the employees. Get them bigger monitors.

u/Wide_Egg_5814 (89 points · permalink)

bro in the back got my posture

u/Fit-Stress3300 (67 points · permalink)

This is exactly how most of the "flashy SF" startups (and major companies) offices look like, including the Chinese engineers.

u/GeorgiaWitness1 (55 points · permalink)

I have a question.

What forbids places like Europe to snap a lot of this guys?

China stands now in a situation of "Elite overproduction" where you should have plenty of this people without space.

Im not saying this guys literally, but overall talent.

Geopolitically is this an issue ?

u/TorturedPoet30 (71 points · permalink)

We don't innovate, we regulate.


Kimi K3 achieves 3rd Place on ArtificalAnalysis, beating out Claude Opus 4.8

846 points · 84 comments · r/LocalLLaMA · by u/MagicZhang

Kimi K3 achieves 3rd Place on ArtificalAnalysis, beating out Claude Opus 4.8

Kimi K3 has achieved third place on Artificial Analysis's Intelligence Index, ranking ahead of Claude Opus 4.8 and demonstrating that open-weight models are approaching frontier-level performance. The model's cost-per-task and output-token efficiency metrics are particularly promising, with K3 using roughly half the tokens of Opus 4.8 for comparable tasks while costing $3/$15 per million tokens—roughly Sonnet-level pricing for Fable-class performance.

Interesting Points
  • Kimi K3 costs $3 input / $15 output per million tokens, placing it near Fable 5 pricing at Sonnet-level costs.
  • Cost-per-task benchmarks show K3 using about half the tokens of Opus 4.8 for comparable work.
  • The model has 2.8 trillion parameters with approximately 1M context window.
  • Weights are scheduled to be released on July 27, enabling third-party hosting and inference competition.
Top Comments

u/ForsookComparison (153 points · permalink)

Waiting on someone to report in after a long session with it. I've seen enough bar-charts for the day. At Sonnet costs and 30 t/s, it better be hyper-efficient at reasoning.

u/bopbop9876 (135 points · permalink)

Here are cost per task and output tokens per task as well. Super promising on both fronts!

https://preview.redd.it/ayxi7od6bndh1.png?width=1753&format=png&auto=webp&s=14190215c0ae612463e1d7e9a7587b2d5e0c5b48

u/LegacyRemaster (33 points · permalink)

https://preview.redd.it/y1o9gzdn9ndh1.png?width=1007&format=png&auto=webp&s=ecf8bcd32522d4397c88647415c2dbfa395394c9

u/CarryAgile3791 (31 points · permalink)

Just wait a few more months, then open weight models will surpass proprietary models. I wonder how Anthropic then will explain their exorbitant prices...

Same story in 1 more subreddit: r/artificial

Kimi K3 landed third on the Intelligence Index, ahead of Opus 4.8, and even GPT-5.6 Sol couldn't take #1 from Fable 5

34 points · 36 comments · r/artificial · by u/hero88645


Anthropic and OpenAI don't have secret sauce

705 points · 269 comments · r/LocalLLaMA · by u/a9udn9u

A discussion about whether Anthropic and OpenAI have any technical advantage beyond scale. The post argues that their moat is simply scale, with rumors that Opus has 5T parameters and Mythos/Fable are 10T parameter models, while open models stayed under 1T for a long time. Only recently was that ceiling broken by DeepSeek V4 and now Kimi K3, with significant performance jumps as parameter size increased. The community largely agrees that the working recipe is known, and the main moat will be compute access rather than technical breakthroughs.

Interesting Points
  • Rumors suggest Opus has 5T parameters and Mythos/Fable are 10T parameter models, while open models stayed under 1T for a long time.
  • The working recipe for LLM advancement is now known, with multiple independent labs reaching similar outcomes.
  • Anthropic had stronger conviction in scaling laws and made the right bets earlier, but other labs are now catching up.
  • The main moat for frontier labs will be access to compute for both training and inference, not technical secrets.
Top Comments

u/stoppableDissolution (394 points · permalink)

They definitely have some very damn good synthetic data pipelines (and resources to run them). Theres no point in more weights if you cant saturate them

u/uutnt (275 points · permalink)

It would seem so. The working recipe seems to be known at this point. The proof is that multiple independent labs have reached similar outcomes. Anthropic just had stronger conviction in the scaling laws, and made the right bets earlier, but other labs are now catching on. I think the labs moat, insofar as they have one, will be access to compute - both for training and inference. Margins will come down though.

u/WaveOfDream (124 points · permalink)

The fact these OSS models can match them while having significant less compute tell us they really don't have any meaningful technical edges, just more resources.

Really most of LLM advancement is scaling and better pipelines/architecture. Not any theoretical breakthrough.


That one IT guy:

635 points · 26 comments · r/ChatGPT · by u/ExpensiveCoat8912

Meme showing an IT professional managing multiple AI tools simultaneously

A meme depicting the modern IT workflow of managing multiple AI tools simultaneously has resonated widely with users. The image shows someone juggling different AI platforms, with comments highlighting how the removal of usage limitations has led to heavy late-night AI usage. Users describe making different AIs argue with each other to reach correct answers as a common workflow pattern.

Top Comments

u/Novel_Style6054 (46 points · permalink)

Lmao this is way too accurate 😭 the 3am AI stack goes crazy.

u/Ok-Hovercraft3823 (22 points · permalink)

i have open ai to thank for the bags under my eyes, they removed the 5 hour limitation and i havent slept since

u/Novel_Style6054 (19 points · permalink)

the real IT workflow is making all 3 AIs argue with each other until they accidentally reach the correct answer 😭

u/Christosconst (7 points · permalink)

All 9am senior software engineers:

u/Straight_Random_2211 (5 points · permalink)

what AI is the black icon? It is not Grok or Gemini icon


you show me kimi k3 is not benchmaxxed, i cancel my claude subscription right now and i go work with open-weights

555 points · 104 comments · r/singularity · by u/98Saman

Meme expressing willingness to switch from Claude to open-weight models if Kimi K3 is proven not benchmaxxed

A meme expressing the sentiment that users would immediately cancel their Claude subscriptions and switch to open-weight models if Kimi K3 could prove it wasn't benchmaxxed. The post has sparked discussion about the value proposition of open-weight models, with some users noting that organizations could achieve Opus 4.8-level intelligence with 100% data security by self-hosting. Others point out that the hardware requirements (GB300 NVL72 racks) make this impractical for 99.99% of organizations.

Top Comments

u/tiger_ace (188 points · permalink)

i don't get this narrative, it's not that hard to just try out different models to get your own feel. everyone should have their own suite of prompts that matter to you and then you can act as the human verifier / benchmark yourself.

benchmarks act as a high-level general heuristic so when you see something like sonnet 5 coming out you know not to expect anything exciting

u/Objective-Picture-72 (61 points · permalink)

The brilliance of KK3 isn't that it will replace your Claude or Codex sub. It's that an organization can have an Opus 4.8 level of intelligence with 100% data security and unlimited ability to customize.

u/No-Head-Royal (37 points · permalink)

A $200 subscription on OpenAI gives you about $14,000 in API use a month, and on Claude, $8,000. Unless you're a power user who ate through that like cake, then I'd advise you to keep your subscription.

That said, the bulk of income for Anthropic is API use, so there is a strong reason to expect significant problems to Anthropic. I wouldn't bet on it happening immediately; institutional inertia is a bitch, but Q3 and Q4 are gonna hurt if Anthropic can't drop Fable 5.1 good enough to stand a generation above Fable 5 and Sol.

u/ToastedandTripping (29 points · permalink)

"The arena includes two hardware platforms, three types of kernels, and four tasks: Attention Residuals and KDA linear attention on an NVIDIA H200 GPU, a 512-head-dimensional MLA kernel implemented from scratch, and a KDA task on a domestically produced GPU. At maximum thinking intensity, Kimi K3's performance is close to Fable-5 (including the fallback mechanism) and significantly outperforms Opus 4.8, GPT-5.6 Sol, and GPT 5.5."

Damn.

https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ

u/Evan_gaming1 (27 points · permalink)

dude thinks open source models are benchmaxxed and closed source models arent im crine


open source models pose EXTREME DANGERS

536 points · 87 comments · r/singularity · by u/Crazyscientist1024

Satirical image about open source model dangers

A satirical post about the 'extreme dangers' of open-source models, playing on the fear-mongering narrative from some AI safety advocates. The post has been interpreted as commentary on how open models pose 'extreme dangers to their bottom line' for closed-source AI companies. Commenters discussed the potential for the US government to outlaw open-weight models, with comparisons to music piracy and predictions that tech-savvy people won't be denied access.

Interesting Points
  • The post is widely interpreted as satire, with commenters noting that open models pose 'extreme dangers to their bottom line' for closed-source AI companies.
  • One commenter observed that 'more intelligence makes more intelligence easier to make. It's an abundant model. Profits thrive on scarcity models.'
  • Commenters predicted the US government might try to outlaw open-weight models, with comparisons to music piracy.
  • The post reflects growing community frustration with the fear-mongering narrative around open-source AI.
Top Comments

u/Ignate (189 points · permalink)

More intelligence makes more intelligence easier to make. It's an abundant model. Profits thrive on scarcity models.

The Singularity means the death of the profit model, not a new profift model.

u/RanklesTheOtter (43 points · permalink)

Extreme dangers to their bottom line, that's what open models pose.

u/jd52wtf (38 points · permalink)

Just wait until they get the US government to outlaw open weight models.

Think it won't happen?


Kimi K3 Benchmarks

503 points · 205 comments · r/singularity · by u/WhyLifeIs4

Screenshot of Kimi K3 benchmark results showing performance across multiple coding benchmarks

A post sharing Kimi K3's benchmark results across multiple coding benchmarks has generated significant discussion. The benchmarks include DeepSWE, Frontier SWE, Terminal Bench 2.1, Program Bench, and SWE Marathon. Users note that K3 achieves Fable-class performance with less than half the parameter size, and at $3 input / $15 output pricing—roughly half the cost of GPT-5.6 Sol and a third of Fable. The results have led to comparisons with the DeepSeek R1 moment, with some users calling it the first Chinese model they would use daily.

Top Comments

u/Eyelbee (141 points · permalink)

Too good. Practically fable class. This will be the first chinese model that I use daily

u/Leading-Shake8020 (124 points · permalink)

https://preview.redd.it/sear4os1pmdh1.jpeg?width=1080&format=pjpg&auto=webp&s=1bf81432ec0b92dcfe5d13c90b0771f3d3cf265a

Coding Bench: DeepSWE, Frontier SWE, Terminal Bench 2.1, Program Bench , SWE Marathon
Source: https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ
Tech Blog Post: https://www.kimi.com/blog/kimi-k3

u/cakes_and_candles (67 points · permalink)

Fable lvl performance with less than half the parameter size, damn

u/Mierzejsky (62 points · permalink)

The White House must be having a real hard time figuring out how to put an export block on something that isn't their product.

u/coinfreekz (54 points · permalink)

And $3 input $15 output, basically half the cost of sol and around a third of fable.


85 more Reddit stories

Updates: 05:30 AM PDT · 07:24 AM PDT · 08:30 AM PDT · 11:30 AM PDT · 02:30 PM PDT · 05:30 PM PDT