· 05:30 PM PDT

Kimi K3 Reshapes the Race as Claude Fable 5 Expands

Overview

China’s Kimi K3 has rapidly ascended the model leaderboards, intensifying the global split between open-weight and closed AI ecosystems while challenging Anthropic and OpenAI’s market dominance. On the product side, Anthropic is rolling out Claude Fable 5 to broader subscription tiers, sparking intense discussion around compute scaling, pricing economics, and DeepSeek’s relentless cost efficiency. Beyond the benchmark wars, the industry is grappling with mounting friction: Satya Nadella calls out data scraping double standards, Linus Torvalds defends AI integration in core codebases, and regulators consider FINRA-style oversight as cultural discourse shifts from capability hype to societal impact. Meanwhile, frontier models continue demonstrating astonishing leaps, with GPT-5.6 cracking decades-old optimization problems and solving international math olympiads without human steering.


Hacker News Stories

GPT-5.6 used a prompt to close a 30-year gap in convex optimization

484 points · 314 comments · by mbustamanter

GPT-5.6 used a prompt to close a 30-year gap in convex optimization

GPT-5.6 Sol Pro successfully derived a quadratic lower bound for deterministic zeroth-order convex optimization, closing a complexity gap that has persisted since 1996. The model generated the core argument in a single 148-minute session using a prompt modeled after OpenAI's recent Cycle Double Cover Conjecture proof, which was subsequently formally verified in Lean. The author argues that while this demonstrates AI's ability to automate mathematical proofs where the underlying techniques are already established, it signals a shift where human researchers must increasingly focus on problems requiring genuinely novel theoretical approaches.

Interesting Points
  • The AI required 148 minutes of uninterrupted computation to produce the main lower-bound argument, despite the author's own year of sporadic attempts and failed runs using GPT-5.4 and GPT-5.5.
  • The mathematical problem concerns oracle complexity in derivative-free convex optimization, specifically addressing whether gradient-free methods fundamentally require order d² evaluations rather than the Ω(d) bound inherited from first-order models.
  • The prompting methodology involved a roughly ten-page document that was partially co-authored by an AI to synthesize closely related literature and explicitly structure the adversarial oracle search.
  • The proof's core construction relies on a family of 'max of affine functions' paired with a specific adversarial querying strategy, closely mirroring techniques from Nemirovsky and Yudin's tight bound for first-order optimization.
  • Formal verification in Lean was completed collaboratively using Sol and Claude, with the author noting that automated formalization is rapidly becoming a community standard for AI-authored mathematical results.
Top Comments

mw67 (13 replies)

Crazy how intelligence is cheap, efficient and commonplace now. We humans better refocusing our energy on our core values/principles, given most of our skills are becoming irrelevant

threethirtytwo (7 replies)

Genuine question: If you still or did think LLMs are just stochastic parrots that just summarize everything and have no form of creativity, what do you think after seeing results like this?

I'm very curious how people reconcile their fear/hatred of AI with actual objective reality. This is actually what interests me most about the whole AI thing. How we tell ourselves what we tell ourselves.

rakel_rakel (5 replies)

I don't think researchers in math/TCS will be made obsolete, but I think it will instead no longer make sense to work on any low-hanging, or even medium-hanging (you know what I mean) fruit. We'll be needed for problems where actual novel approaches are needed.

I wonder how this compares to what we see happening with "juniors" in software development? In math research, do you also get the training for the profession from working on the low hanging fruits for a while, to then move to the medium-hanging, and later go on to work on previously unsolved stuff?

applfanboysbgon (4 replies)

Two points:

  • Hasn't been peer reviewed yet, so take with a grain of salt. This applies to all claimed proofs, not just AI-generated ones. Even humans hallucinate proofs too!

  • The prompt is on page 27 here[1]. It is ten pages of advanced mathematics priming the model in the right direction, apparently informed by a year of prior research. That doesn't invalidate the result if it is genuine, but it is worth noting that this wasn't a matter of "ChatGPT, solve this unsolved problem. Make no mistakes." and required substantial domain expertise and human research beforehand.

[1] https://arxiv.org/pdf/2607.13335

alternator (2 replies)

I know a bit about this field. This conjecture reads as somewhat more niche than the cyclic double cover conjecture recently proved by OpenAI, but nevertheless represents a real contribution.

You want to know how long it takes to solve an optimization problem, in this case over convex, lipschitz functions. (The restriction to a spherical domain is not really a restriction, you can just change variables for any bounded domain.) Anyway, showing upper bounds on time complexity is "easy" because it's just the runtime of your algorithm. Showing (nontrivial) lower bounds is usually much harder because it requires constraining all algorithms.

This proof apparently shows that the lower bound time complexity is equal to the time complexity of an existing 30-year old algorithm: it requires Omega(d^2) function evaluations to solve over this class of functions.

My gut says likely implies that d is the minimal number of evaluations if you have a gradient oracle because you can approximate a gradient with d function evaluations, but I'm not sure how hard it is to make that rigorous.


Why do AI company logos look like buttholes?

411 points · 139 comments · by miniBill

Comparison of AI company logos arranged in rows

An essay examining the striking visual uniformity among AI company logos, noting that a disproportionate number share a circular template with a central opening, gradient, and radiating curves. The author attributes this trend to corporate design-by-committee processes, risk-averse stakeholders averaging out creative suggestions, and the psychological appeal of circles as symbols of wholeness. Only DeepSeek and Midjourney break the mold, incorporating sea-related themes instead. The piece concludes that while this aesthetic currently builds trust, it stifles creative differentiation and mirrors broader cycles of tech design conformity.

Interesting Points
  • FastCompany identified the trend in 2023 but published it under the title 'The AI boom is creating a new logo trend: the swirling hexagon' to avoid direct anatomical references.
  • Anthropic's Claude logo bears a striking resemblance to an explicit anatomical drawing from Kurt Vonnegut's 1973 novel Breakfast of Champions.
  • OpenAI's official brand guidelines justify its circular logo by citing 'the fluidity and warmth of human-centered thinking through the use of circles as a core design principle.'
  • The article traces tech branding conformity cycles: 1990s-2000s 3D gloss, 2010-2013 skeuomorphism, 2013-2018 flat design, and 2018-2022 neomorphism.
  • Among major AI firms, only DeepSeek and Midjourney break the circular central-opening mold, and both explicitly incorporate sea-related themes into their branding.
Top Comments

SkyMarshal (10 replies)

Claude is the only one that looks like an asshole. The rest are just circular, or not even that. Does every circle in the world look like an asshole? Car wheels? Pizzas? Camera Lenses? Ferris wheels? This is like a Rorschach test.

VladVladikoff (5 replies)

Then came the redesign: a perfect circle with a subtle gradient and central void.

I don't see the gradient, their logo is black and white. Where's the gradient? Was this written by an AI hallucinating?

sparsesignal (5 replies)

Have you tried clicking the Claude logo? https://x.com/ertug/status/2072339797708849398

designerarvid (3 replies)

They're apertures; symbolically things emerge from them.

"…but that's really all a butthole is, an aperture"

  • Louis CK

anal_reactor (1 reply)

It's less about how most of them look like different stages of fisting party, and more about the fact that they have zero distinctive elements. Despite the amount of ads being shoved down my asshole, I still find it difficult to recognize different brands. Meta and Gemini being the worst offenders, if you show me these two logos in a week and ask what they are, I'd have zero clue they're even company logos. This is the level of failure we're talking about.

I think this is peak millennialism because the design is supposed to be as inoffensive as possible and therefore shows absolutely nothing.


What AI did to stackoverflow in a graph

354 points · 417 comments · by secretslol

A visualization of Stack Overflow's question volume over time shows a gradual decline from 2016 onwards, with a sharp acceleration after ChatGPT's release in late 2022. The number of questions halved from approximately 100K per month in November 2022 to 50K one year later, before continuing to fall toward near-zero levels. Commenters debate whether AI was the primary cause or merely accelerated a decline that had already begun due to SO's toxic community culture, hostile moderation policies, and the rise of alternative platforms like Reddit and GitHub issues.

Interesting Points
  • The number of questions on Stack Overflow halved from roughly 100K per month in November 2022 to 50K exactly one year later.
  • The decline began gradually around 2016-2018, well before ChatGPT, suggesting SO was already losing relevance to alternative platforms.
  • Commenters note that SO's decline was also driven by its own policies: closing questions as duplicates of outdated answers, hostile moderation, and sunsetting features like the Jobs board.
  • One commenter observed that their company's internal Slack help channel has become almost entirely inactive, with questions that could have been solved with a single LLM search replacing peer-to-peer communication.
Top Comments

lynndotpy (33 replies)

Any social organization needs to carefully consider their inclusion-exclusion curve with intentionality.

I think a lot of people might balk at the word "inclusivity" today, but StackExchange had ridiculously high barriers to participation, making it inclusive to the long-time users on the site, but exclusive to the newbie participants who found themselves blocked for asking questions. They slowly killed the site in this manner.

The community might have survived this folly, even with AI, because it was still the best place for people with qualms about AI to ask questions... Except until StackOverflow management alienated those users, too, by shoving AI down their throats in every facet of the site.

Even I had internalized the vagaries and neuroses of the SO community but I had heavy reticence to ask questions, knowing I'd have to consider all the ways a bully eager to use their powers might misunderstand me. I can't imagine asking a question there without having had lurked for longer than a typical Bachelor's + Masters program.

Peak at 207K, minimum at 588. That might be an incomplete date point, so using the next most recent value 1226, StackOverflow has lost 99.41% of its activity.

dfabulich (8 replies)

The graph proves the cause of their decline was AI, and not aggressive question moderation.

Ask yourself: in what year did it become difficult to ask questions on Stack Overflow? 2014? 2016? 2018? 2020? Aggressive question-closing was part of their design from the very beginning. Their high barriers to question-asking was the cause of their rise, as their primary user was never question writers: it was Google, and anonymous Google users. The whole thing was an SEO play from start to finish.

It's fun to imagine that their aggressive moderation was the "real" cause of their decline. It feels so gratifying, doesn't it? Finally those assholes got their comeuppance, because of their bad behavior!

But that's not why they failed. They failed because SEO businesses can't survive when AI answers the question directly, without referring you any traffic.

(The same thing is happening to Wikipedia, BTW, which is also aggressively moderated.)

bartread (2 replies)

I just noticed that by around 2015 it had got super toxic and snippy.

Usually I'd find answers on SO. Relatively rarely I'd ask questions but, when I did, I'd always try and follow the netiquette rules of yore, and think in terms of, if I was a support engineer trying to help with this, what would I need to know?

Because I have supported products, and we've all seen enough bug reports and questions come in that we can tell when someone is going to be easy to help - even if they have a particularly tricky problem - versus someone who's going to prove more challenging.

So I had this question about Elasticsearch, and it was at a time when the documentation wasn't great, and you were actively encouraged to go on SO and tag your question to get help.

I wrote out in detail what I'd done, where I'd got stuck, what I'd read and tried to get unstuck, etc. It probably took me 30 minutes or more to pull everything together into a coherent post.

The very first comment was from some insufferable bellend saying, "Oh, so you want us to do your work for you, are you going to pay us too?" or words very much to that effect.

Literally, WTF? Why even post that? If you don't want to help the option to simply go away without getting involved is always available.

IIRC I didn't actually end up finding a solution via SO and instead layered some godawful hack on top of Elasticsearch to get what we needed - because I simply had other work to move on to and I'd already spent a lot of time on the problem.

But I think that was the last question I posted on SO, and maybe the last time I posted anything on the site.

As the years wore on I simply started finding it less and less useful, with often incorrect answers marked as accepted and - if you were lucky - the correct answers marked might be buried further down.

And then there's what they wanted to charge for job ads versus how effective those ads actually were - again, this was better in their earlier years.

SO started out well - genuinely a breath of fresh air - but as time went on it felt like they thought their model was the last word in online help forums and they didn't want to evolve to address its flaws, even if that had just been dealing with the toxicity, and the karma farming.

And so this is the result - a site that, like the dinosaur in A Sound of Thunder, is dead but perhaps hasn't realised it yet - and, at this point, the way I feel is simply good riddance. It's a shame, but - as you said - they did it to themselves.

DrewADesign (4 replies)

Their rules, (I believe unintentionally) give iron-fisted fiefdom rulers a toolbox of justifications to control and alienate under the guise of protecting the quality of the site data. I honestly don't even think most of the control freak mods could objectively judge the propriety of their actions because it was all encouraged by the rules. (And I do not think this was universal among the mods, but it was certainly endemic to the site culture.) Well, the outcome was predictable.

Before I worked as a web developer, I was a formally educated and credentialed professional in a non-computer-related field with a pretty high barrier to professional practice, but a lot of passionate hobbyists. When I found the related low-ish volume SE, I excitedly poured hours into writing authoritative, well-informed, well-cited, thoughtfully worded, and concise but layperson-friendly answers. I also provided encouraging and positive, but usefully critical feedback to people that missed the mark. I knew how negative the format could be after using SO for years, so I bent over backwards to avoid discouraging newcomers with a punitive or imperious tone. People seemed to find my contributions useful because I became the top contributor in something like two weeks, and still regularly get points for things I wrote over a decade ago.

Some mod— a hobbyist with far less knowledge and experience, but a serious case of Dunning-Krueger— probably got annoyed that I was getting more votes than them because one day they started nitpicking the hell out of every goddamned word I wrote. I pretty quickly got fed up, and stopped participating about a month after I started.

::slow clap:: Well they might not have protected the utility or integrity of their knowledge base, but they sure protected the integrity of a bunch of people's egos. That's something, right?!

godwinson__4-8 (1 reply)

In the age of gratuitously encouraging LLMs and constant slop I sometimes miss the way in which stack exchange punished people for putting in low effort.

I recently asked a LLM a question simply out of newfound habit. I realized while reading the line that this command seemed very familiar. I had bookmarked the exact same command in stack exchange many years ago. Of course the fact that stack exchange traffic has drastically declined came into my head. As well as the fact that this was predicated on them doing all that work so the LLM could essentially scrape, steal and now serve to me in this (for now) more convenient form.

These places are ultimately transactional. Taking personal offense to someone insulting your question is something you should learn to get over. The site overall would suffer. The vast majority of traffic wasn't people asking or answering questions it was using it as a search engine.

In Google I used to put in site:stackexchange or w/e because I knew answers there were less likely to be dead ends. Yes this is because some people got their feelings hurt. Iirc there may have been specific controversies re toxicity in meta, but scoping this purely to the question and answering side I never felt my own experience on stackexchange reflected something personal or done purely out of spite. I saw it as a way to ensure the site remained valuable as a place for high quality answers. Ensuring people have to write better questions is part of that. I often wish my LLM was meaner. My first stack exchange ribbing I just learned to ask questions better. I felt a human annoyed at my admitted laziness. This was valuable feedback. Should I have went boohoo and blamed them instead?

Now we have people wanting to fuck their AI girlfriends and waste your time with their LLM authored blog. If you think the problem with stack exchange was your feelings got hurt I just worry about what road you are leading us down. Having to earn your way in is part of human society. A LLM that tries to make you feel like a genius so you keep coming back is alien to most of how humans got here. I think the decline of stack exchange is less about a change in traffic patterns and more reflective of a continuing change in culture. I'm also guilty of course. But when I was reminded of my old bookmark, I went back and browsed other old stackexchange bookmarks. I did miss the very human nature of these questions and answers and comments. Yes including the occasional scolding and bickering. Sort of like this site. Filled with humans, yet not a complete cesspool. Oddly hopeful in our present age.


Setting up your spare Mac for Claude Code to control, a step-by-step guide

169 points · 127 comments · by ykev

This guide walks through transforming a spare Mac into a dedicated, always-on environment for Anthropic's Claude Code agent, isolating potentially risky autonomous operations from your primary workstation. The author recommends running the agent on a freshly wiped machine with a local account that lacks an Apple ID, then controlling it remotely via SSH from a main Mac or the Claude mobile app. By configuring passwordless sudo, persistent wake settings, and a custom tmux LaunchAgent, users can safely enable Claude Code's computer use capabilities and browser automation.

Interesting Points
  • Containers like safeclaw still route network traffic through the host machine and lack native support for Mac-only applications like Unity.
  • macOS ties screen capture and input permissions to the active GUI session, requiring a persistent tmux server with an anchor session to bridge SSH processes to the display.
  • The guide provides a custom ic shell wrapper that manages multiple concurrent claude sessions, tracks their states, and supports a /remote-control mode for driving the agent from a smartphone.
  • Tailscale can extend LAN-only control to any network by routing traffic through encrypted WireGuard tunnels while keeping MagicDNS hostname resolution functional without the .local suffix.
Top Comments

catoc (14 replies)

I just cannot come up with a good AI-is-actually-24/7-helping-me-out use case.

Please help: I want to need this!

deadbabe (8 replies)

I still don't understand what these freaks are doing running these agents 24/7 on machines. What are they doing? Managing a todo list? You mean crossing items off as you complete them? Research tasks? To do what?

Never really get good answers. There is no killer app. Just bikeshedding.

ianm218 (2 replies)

Many Claude Code power users don't really use IDEs anymore, so the only purpose of them working from their laptop instead of a phone is because that is the normal way to do it.

Here is a real use case: you are are responsible for some alerting channel. You have datadog/ cloud logging/ github all connected. You see a bunch of alerts come through while you are out and about and you prompt CC to investigate - Claude triages and says "all of the sudden you are getting time outs from this bank API your company partners with, this started an hour ago. It's happening on ~15% of requests". So you ping the guy at your company who does vendor relationships and go back to your weekend.

This is a non hypothetical example. Obviously it would be better if your job had a real on call rotation and more robust alerting and you wouldn't be getting slack alerts on the weekend… but I take the approach this job affords me a lot of nice flexibility so it's ok

esaym (5 replies)

Outside of the article's mentioned graphics development, there is no reason to isolate an agent using actual hardware. I threw together this script[0] using libvirt to give claude its own graphical desktop env to be able to do user acceptance testing with Chrome. It has full root and can do what ever. If it makes a mess, I can dump and reinstall in seconds.

0: https://gist.github.com/smith153/04b4068b5a2d7b234f1c3d5992dafe25

addajones (4 replies)

I currently have Claude Desktop installed on a separate Mac mini M4 and control it with Dispatch. Is there a reason to do this method, it still seems the way I have it setup it has full control over the local account I gave it on the Mac mini.


Mayor Mamdani Says Landlords Can't Use AI Images to Advertise

119 points · 48 comments · by gnabgib

New York City Mayor Zohran Mamdani has released a Rental Ripoff Report recommending that landlords and real estate agents disclose when rental listings use AI-generated or AI-edited images. The proposal emerged from public hearings across all five boroughs where thousands of residents shared complaints about deceptive practices. The initiative covers both fully AI-generated photos and digital alterations, with particular concern for tenants signing leases remotely when physical conditions contradict advertised imagery. The disclosure mandate is bundled with broader tenant protection reforms including code enforcement modernization and formalized tenant union recognition.

Interesting Points
  • The disclosure policy was developed after the mayor met with 2,400 New Yorkers across each borough during his first week in office.
  • The administration noted that tenants signing leases remotely are especially vulnerable when physical conditions contradict advertised imagery.
  • This mandate is bundled with broader reforms aimed at modernizing the city's code enforcement systems and formalizing tenant union recognition.
  • The announcement follows a separate click-to-cancel rule Mamdani introduced the previous day targeting subscription-based companies like Adobe.
Top Comments

sssilver (7 replies)

Isn't literally every photograph taken with a modern iPhone technically an "AI-generated / AI-edited image"?

mingus88 (2 replies)

no, the photos you take with the lenses on your phone are not AI generated. They are generated from the sensors on your phone.

Have you seen some of these listings? We are talking about retaining walls invented where they can't exist, work displayed that hasn't occurred, etc. if you show up to a property and it's materially different than the picture that got you there, that should be illegal.

If you want to make an argument that "everything is AI now" go for it. But I'm happy to see existing false advertising laws evolve as technology evolves

avaer (2 replies)

There's several other areas that would be good to categorically ban AI usage from:

  • gambling
    • dating
    • hiring
    • advertising

It shouldn't even be controversial that this would be broadly good for society.

I say that as an AI maximalist: I fully trust AI with these things. I do not trust the humans using the AI.

plants (1 reply)

This is awesome! StreetEasy is how many New Yorkers find apartments. In the past few years, it has been flooded with AI-staged apartments. The AI stagings warp the room to fit furniture that would 100% certainly not fit there. It's deceptive, and I'm glad it at least requires disclosure now (although I wish it were fully banned)

DangitBobby (1 reply)

HN title is missing the operative word "secretly". The real title:

Mayor Mamdani Says Landlords Can't Secretly Use AI Images to Advertise Properties

The article contents align with the real title: you just disclose AI usage when advertising rentals.


A grumpy screed about AI in software engineering

60 points · 77 comments · by ssutch3

Software engineer Samuel Sutch argues that AI-generated content has become ubiquitous and non-negotiable in the industry, effectively transforming daily development into a frustrating, joyless slog. He notes that AI now comprises nearly all pull requests, design documents, and even leadership roadmaps, leaving little room for traditional craft. While acknowledging that opting out is no longer professionally viable due to hiring practices and market saturation, he suggests that non-AI development may only survive in isolated niches rather than achieving commercial scale.

Interesting Points
  • AI content now accounts for nearly 100% of code, PR descriptions, and design specs in many organizations, with leadership roadmaps reportedly at least 50% AI-generated.
  • Job interviews have shifted to using candidates' opinions on AI as a primary filter, effectively screening out those who refuse to adopt the tools.
  • The author, with two decades of engineering experience, describes the shift as making software development an 'absolute slog' compared to the creative satisfaction it once provided.
  • Independent, non-AI software development faces steep commercial hurdles due to market saturation with AI-generated applications, a trend the author claims extends to major players like Apple.
  • The author maintains his personal website and blog are entirely hand-written, using AI strictly as a work requirement rather than a personal preference.
Top Comments

rDr4g0n (0 replies)

"software engineer" covers a wide range of responsibilities. some engineers spend all their time contributing to the physics in a massive graphics engine. some spend most of their time as the sole maintainer of a web api that mostly plumbs together a bunch of services and provide data to UIs. some are frequently spitting out new services while other are maintaining 10 year old legacy codebases. the work we need to do, the tradeoffs we need to make, its just so vast.

now couple that with the wide range of variety (and dysfunction) that come from the business and organization. being a dev for a company which sells software is a whole different beast than being an in-house dev for a company which sells other products. startup, big org, old org, new org, strike team, the list goes on.

finally, engineers enjoy different aspects of building software more than others. the puzzle, the crafting, the api, the data structure, the architecture, solving end user needs, so many facets to what makes it enjoyable as a career.

this is a big reason why ai discourse online is so uneven and often unhelpful. were all assuming too much about what being a software engineer is; forgetting that my day-to-day can be wildly different than yours.

a lot of ai speedup comes from effective delegation of work. i do it, but i don't like it. it actively makes me feel tired in a way i don't feel when i do the work myself, or when i mentor and grow my team to do the work.

many other devs don't feel the way i do. some suddenly have these ai tools that make them feel more empowered and energized and excited.

llm's are genuinely a new computing paradigm because of their nature: plain human language as the interface. this is different enough that mapping our past computing concepts (like the evolution from assembly to C) may not hold as well as seem they should.

this all digs deep into who we are as individuals with specific desires, career history, goals, a specific role at a specific company. the part that's fucking beautiful and unique and human.

But if the aggressive industry-wide layoffs aren't enough of a clue, these giant ass companies (and the ai gatekeepers driving this shift) are not compatible with individual humans.

i dont mind llms as a tool, but im done with the perverse level of greed and inhumanity they seem to inspire in corporate leaders.

stiglitz (0 replies)

It's a Skinner box of productivity. Of course it's polarizing.

My own enthusiasm varies wildly day to day. I've been sick of squinting at punctuation marks and opening sequences of of tabs to trace through annoying abstraction layers for a long time; so I'm really excited to see those activities effectively automated away. But the same tools enable some really bad behavior (in myself and others), which really gets me down.

dd8601fn (0 replies)

It's the NES Game Genie.

Unlimited lives and the ability to walk through walls is great fun for a bit. But also, you never actually played the game and it kinda ruined it for you, for all time.

You've seen the outcome, you solved nothing, learned nothing, and there's zero reason for you to ever be proud of it.

The outcome isn't yours, it's the chatbots. Maybe you lie to yourself a bit… but you were barely necessary.

shric (1 reply)

"Wholly" was too strong a word now that I think about it.

But I'll give two main things I find unpleasant.

One: Normally I like working with other people. There is nothing more satisfying than collaborating on a complex problem with smart and conscientious colleagues. However, the bad experiences are now amplified:

  • a random colleague gets sent down a wild gooose chase of an LLM claiming there is an issue with something I wrote. It's a full blown hallucination because it found a confluence page which mentioned me and just decided to connect things together.

  • another colleague is suddenly empowered with being able to vibe code websites and is now copying and pasting my replies into an LLM and pasting the LLM back at me over Slack.

  • a colleague gives a quick drive by vibe coded PR to add a feature into something I own. It reimplements the same business logic that is already done elsewhere. It takes me more energy to explain all this than just do it myself.

There are dozens more examples. In isolation I can shrug each of these off, but it's draining. Large organisations have become more "agile" with local optimisations and missing the bigger picture.

Two: while I agree that everything is getting done faster and I do get satisfaction out of that, I'm a deep in the weeds technical person. I don't have it in me to be a big picture person or a people leader. I love having to think about algorithms, data structures, schemas, etc. While I can use AI has a pair programmer, it's doing a lot of the thinking for me and it's usually faster. So yes it makes me a better engineer, no I don't enjoy not thinking as much. I don't enjoy shifting solely to a higher level of abstraction. Maybe that makes me a bad engineer, but I cannot change how I feel about it.


AI Mania Is Eviscerating Global Decision-Making

43 points · 6 comments · by bertman

AI Mania Is Eviscerating Global Decision-Making

A technical consultant who has engaged with hundreds of professionals globally argues that widespread corporate AI adoption is driven by political coercion and hype rather than actual productivity gains. Organizations are experiencing decision-making paralysis as executives and employees feel compelled to profess unwavering belief in AI's transformative power to keep their jobs, while skepticism is actively punished. The author reports a 0% observed success rate in AI projects over a 1.5-year period, with companies resorting to AI-washing initiatives and gaming metrics to satisfy leadership mandates. Ultimately, the article contends that this coordinated corporate mania is destroying rational strategic planning and forcing competent professionals into survival mode.

Interesting Points
  • The author reports a 0% success rate across all observed AI projects over a 1.5-year period, noting that internal chatbots frequently see no uptake because companies lack quality documentation for the models to reference.
  • Customer-facing AI implementations often fail catastrophically in practice, exemplified by a Mitsubishi automotive support bot that promised a callback for a six-month-old issue but never followed up, effectively hiding the failure from internal metrics.
  • Employees are being evaluated on gameable AI metrics like token leaderboards, prompting engineers to write semi-plausible prompting loops just to consume tokens while watching Netflix to avoid being flagged for low usage.
  • Funding and hiring requests now routinely require staff to demonstrate they attempted to solve problems with AI first, creating a bureaucratic drag where failure to comply results in automatic denial or termination.

44 more Hacker News stories

Reddit Stories

I accidentally started a new chat

1057 points · 663 comments · r/ChatGPT · by u/Natural-Stomach-7797

I accidentally started a new chat

A humorous meme post about accidentally starting a new ChatGPT chat, which resonated widely with users who relate to the frustration of losing conversation context. The post became one of the most upvoted in the subreddit's history, reflecting the community's shared experience with ChatGPT's chat management system.


Tibo Strikes Back.

683 points · 65 comments · r/OpenAI · by u/PranayJhaTheMan

Tibo Strikes Back.

A post about Tibo (likely referring to a notable figure in the OpenAI/ChatGPT community) responding to criticism or controversy. The post generated significant discussion in the OpenAI subreddit.


What kind of dark magic is Deepseek using?

557 points · 161 comments · r/LocalLLaMA · by u/Fuckinglivemealone

What kind of dark magic is Deepseek using?

A discussion about DeepSeek's remarkable cost efficiency and cache hit rates, with users sharing examples of harnesses that process millions of tokens for under a dollar. The conversation explores DeepSeek's technical optimizations including hybrid attention mechanisms (CSA compressed tokens with selectors and HCA dense attention with compressed tokens), and how their priority on optimization sets them apart from US competitors like Anthropic and OpenAI.

Interesting Points
  • Users report harnesses processing millions of tokens for under $1, with one example showing 9 million tokens for approximately $0.90.
  • DeepSeek uses hybrid attention: CSA (compressed tokens with a selector that selects most important memories) and HCA (normal dense attention with heavily compressed tokens), with the most recent tokens using dense attention without compression.
  • On OpenRouter, DeepSeek's pricing is about 2x cheaper than the next cheapest option for the same model.
  • Contributors note that cheaper overall costs from workers, electricity, chips, and government support also contribute to the low pricing.
Top Comments

u/CalamityMetal (291 points · permalink)

They are known for their cache hit rate on an incredible scale. There are harnesses that fully make use of it like Reasonix. There's numerous posts talking about it, where people use like 9mil tokens for like $0.90 or something

u/shy_monkee (104 points · permalink)

It's not subsidisation, because the other providers offer similar prices for the same model. They are just so so good at optimisation, because it's a priority for them, unlike Anthropic or OpenAI.

u/Nicking0413 (39 points · permalink)

Hybrid attention. One is CSA, which is compressed tokens with a selector that selects most important memories (compressed tokens), and HCA, which is normal dense attention with really heavily compressed tokens. On top of all that, the most recent tokens use dense attention with no compression. I think bycloud has a video on this.

Found it, it's this one. https://youtu.be/gC76aeibdFA?si=GZd8y8A0rDlUr5YL

I got all of my information from this so let me know if he's wrong. Also, I think the cheaper overall price (workers, electricity, chips, etc) and government support/control also contributed to the low price


Bad vibes from the "Head of Strategic Futures at OpenAI"(X: @deanwball)

447 points · 237 comments · r/singularity · by u/jvnpromisedland

Bad vibes from the "Head of Strategic Futures at OpenAI"(X: @deanwball)

A post criticizing Dean Ball, OpenAI's head of strategic futures, for his recent Twitter posts arguing that open-weight models are a dystopian threat and that AI should remain a commoditized good controlled by closed labs. The community has largely pushed back, calling his arguments ideological blathering devoid of reasoning and pointing out the irony of an OpenAI executive opposing open-source AI.

Top Comments

u/DownHatter (488 points · permalink)

If a future dominated by open weight models is a dystopian hellscape, then I don't even want to imagine one dominated by closed models owned by a few.

u/fingertipoffun (159 points · permalink)

Guy selling Oreos gets angry about people getting free cookies. Calls getting free cookies a dystopia.

u/riscum (152 points · permalink)

How the f is open source the communist, regime controlled? Literally their models need their super duper capitalist government to allow it to go to market.

These people are brain washed.

u/Background-Wafer-548 (135 points · permalink)

Sure, and being at the mercy of an oligopoly of closed model frontier labs engaging in obscene wealth concentration isn't tantamount to a "dystopian hellscape".

Whine some more.

u/ThePaSch (69 points · permalink)

This guy's entire Twitter account is nothing but ideological blathering devoid of any reasoning or even attempts at arguing at why his positions are supposedly the best and most coherent way to proceed. This excerpt is essentially a microcosm proof of his entire MO: AI as a commoditized good is "a dystopian hellscape" and proponents of it "don't know what they're doing" with not a single argument for why or how, with a transparently technofeudalistic "bad for business" closer that does nothing to reinforce any of the nonsensical waffle that precedes it.

The funniest thing is just how woefully thin-skinned the guy is. He'll eagerly fire off butthurt subtweets the moment any of his nonsense meets any coherent and substantiated rebuttal like some teenage debate club president with more ego than wit.

I'd say he should just be ignored, just as much as any of the other "policy experts" or "strategic heads" of the very labs who the rapid progress in open-weights modeld is slowly stripping of any justification to exist.

Same story in 1 more subreddit: r/OpenAI

OpenAI's head of strategic futures thinks open-weight models are a 'dystopian hellscape'

36 points · r/OpenAI


Think this could happen to OpenAI / Anthropic?

438 points · 120 comments · r/ChatGPT · by u/PsychologicalBox5208

Meme image depicting financial collapse scenario

A meme comparing the trajectory of AI companies to a financial collapse scenario has sparked extensive discussion about the sustainability of the current AI investment bubble. Commenters debate whether one of the major AI companies will face deep financial trouble within five years, with some pointing out that the market's continued confidence despite mounting evidence of unsustainable spending is itself a sign of the bubble. Others argue that the companies developing ASI will ultimately control the global economy, making current financial concerns moot.

Interesting Points
  • One commenter predicts that within five years at least one major AI company will be in deep financial trouble or gone entirely, with the market gradually questioning where money actually goes
  • Another commenter notes that whoever develops ASI will swallow the entire global economy and harvest every atom in the solar system, which a third commenter calls the exact kind of pitch that voodoo-investors evaluate at ridiculous amounts
Top Comments

u/mca1169 (108 points · permalink)

within 5 years at the very least one of the major AI companies will be in deep financial trouble or gone. doubt is already creeping into the market gradually. by this time next year everything will be noticeably different with a lot more questioning where the money actually goes and what is sustaining the bubble. the fact that it hasn't popped yet is crazy and shows just how determined these companies are to fool people into thinking everything is fine.

u/CriticalDiscipline4 (40 points · permalink)

Whoever develops ASI will swallow the entire global economy, and then harvest every atom in the solar system. You people aren't thinking big enough.

u/rakuu (12 points · permalink)

Nobody's running Kimi K3 or other theoretical frontier open models on local hardware, unless you're a billionaire.


built a tool that hides secret messages inside innocent-looking LLM conversations

384 points · 48 comments · r/ChatGPT · by u/Nethical69

built a tool that hides secret messages inside innocent-looking LLM conversations

A developer has built a tool called conversation-steganography that hides secret messages inside seemingly normal LLM conversations. The tool uses local model configuration to encode covert messages within the text of chat exchanges, allowing two parties to communicate secretly through what appear to be innocuous AI conversations. The repo is available at github.com/nethical6/conversation-steganography.

Interesting Points
  • The tool encodes covert messages within the text of chat exchanges using local model configuration.
  • It allows two parties to communicate secretly through what appear to be innocuous AI conversations.
  • The author notes the tool is more interesting from an information theory perspective than practical.
Top Comments

u/Nethical69 (84 points · permalink)

repo: https://github.com/nethical6/conversation-steganography

u/madsci (74 points · permalink)

Very cool. I don't know about practical but that doesn't really matter - it's neat from an information theory perspective. I think the next step would be having it adhere more to a specific overt text. It'd reduce the density of your covert encoding and you'd need a lot more text to convey the same message but it wouldn't stand out as much.

u/puppy_yuppie (20 points · permalink)

So how does the other user get the same config for their local model to be able to send messages back and forth? And what prevents a third party from getting the same config? Just curious. Cool idea though. With privacy continuing to be attacked, I'm all for this

u/Vas1le (29 points · permalink)

Nsa ain't gonna like that


China's Xi Jinping Wants AI to Be Open to the World—and Out of America's Control

370 points · 256 comments · r/ArtificialInteligence · by u/aacool

Chinese President Xi Jinping has called for AI to be open to the world and free from American control, marking a significant rhetorical shift in China's approach to the technology. The post has sparked discussion about the irony of a country with heavy domestic internet censorship advocating for open AI, and whether China's open-weight model strategy is genuine or a byproduct of US export controls limiting their compute access.

Top Comments

u/zoratosthenes (111 points · permalink)

Good for the world

u/Felfedezni (33 points · permalink)

Totally agree with Xi.

u/Kingalec1 (23 points · permalink)

Said the country with the most controlled of their tech . Likewise , I agree with the statement however I disagree with the merit to achieve true AGI .

u/orph_reup (12 points · permalink)

Damn, the copium of western chauvinists on this topic is better than expected.

gif

u/streetscraper (10 points · permalink)

Dude doesn't even let his own people share photos of cats freely and people here believe he wants software to flow freely.

Same story in 1 more subreddit: r/artificial

Xi Jinping calls for more open-source AI: 'China is ready to be more open'

235 points · 98 comments · r/artificial · by u/esporx


Kimi moment. I think the writing is on the wall for Anthropic and OpenAi

345 points · 125 comments · r/LocalLLaMA · by u/Difficult-Top9010

A discussion about Kimi K3's emergence as a competitive open-weight model that challenges the dominance of Anthropic and OpenAI. The poster argues that the accelerating progress of Chinese open-source models (Kimi, Minimax 3 Pro, GLM 5.3) demonstrates the strength of open source, while enterprises may increasingly distrust US-based closed models. Commenters debate whether the benchmarks reflect genuine capability or overfitting, with some sharing personal testing experiences.

Interesting Points
  • The post highlights that Minimax 3 Pro (2.7T parameters) and GLM 5.3 are also accelerating progress alongside Kimi K3, demonstrating the breadth of Chinese open-source development.
  • One commenter tested Kimi K3 on electronics (a field not covered by benchmarks) and found it gave a more superficial answer than Fable 5 while being more expensive, with CoT that 'basically boils down to wasting tokens on repetitions and loops.'
  • Another commenter noted that Kimi K3 could potentially become the primary source for agentic distillation of synthetic data, which would improve all downstream models.
  • A commenter pointed out that Anthropic and OpenAI are more than just models—they include supporting tools and ecosystems, and open-weight alternatives provide needed competition to keep costs in check.
Top Comments

u/iam_maxinne (215 points · permalink)

I think you are happy a little bit too early… Let's celebrate when the model reach HF, and some rich fellas here validate the results…

u/triynizzles1 (101 points · permalink)

I think kimi is good enough where the majority of synthetic data can be agenticly distilled from k3 instead of open AI and anthropic. Which is oddly good news for both of them.

u/g_rich (89 points · permalink)

I think it's important to realize that Anthropic and OpenAI are more than just their models. They are the sum of all their parts which include the models, supporting first party tools and the ecosystems built around them.

The release of open weight models such as DeepSeek, GLM and Kimi that compete directly with the likes of Anthropic and OpenAI is important not because they are better than their rivals but because they provide the competition needed to keep costs in check while also providing a clear alternative.

u/Mysterious-Duty2101 (75 points · permalink)

I was pretty skeptical about the benchmarks, so I asked a question in the field of electronics, which is an area that almost no benchmark takes into account. The question was about the behavior of a specific type of transformer under a given condition. It's not obvious and requires the model to have a good understanding of magnetic fields.

Fable 5, with thinking set to medium, cost $0.06, took a few seconds to think, and gave a detailed answer.
Kimi-k3 cost $0.07, spent minutes thinking, answered more superficially, and ended up being more expensive. The real disappointment was looking at its CoT, which basically boils down to wasting tokens on repetitions and loops.

And once again, the comments in this subreddit seem to be just a bunch of people commenting without even having tested the models. Just one question outside the area of web development is enough to reveal which models are actually smart, and not just parrots repeating stupid JavaScript code.

EDIT:
Just to be clear: Kimi-k3's answer wasn't bad. My main issue is really just the cost. What I said in the previous paragraph wasn't aimed at K3.


Claude on X: Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to

308 points · 60 comments · r/singularity · by u/vronikas

Anthropic announced that Claude Fable 5 will be included in all Max and Team Premium plans starting July 20, at 50% of current usage limits. Pro and Team Standard users will continue to access Fable via usage credits and will receive a one-time $100 credit. The announcement reflects Anthropic's struggle to balance surging demand for Fable 5 with compute capacity constraints. Users discuss whether this expansion will be sustainable given the model's resource intensity and upcoming cheaper variants.

Interesting Points
  • Fable 5 will be included at 50% of current limits in Max and Team Premium plans, suggesting Anthropic is managing compute demand carefully.
  • Pro and Team Standard users receive a one-time $100 credit for Fable access.
  • One commenter noted that Claude Mythos Preview was initially priced at $25/$125 per million tokens, compared to current API pricing of $10/$50, suggesting Anthropic has a history of reducing prices after initial launches.
  • Users in healthcare report that Fable flags every single prompt related to biology, making it unusable for their work despite its capabilities.
Top Comments

u/Beatboxamateur (74 points · permalink)

I don't think this is gonna last long, with a more capable Fable model at a reduced cost coming in the next month or so.

If people don't believe me on the reduced price aspect, just look at what the initial price for Mythos was when the model was announced for Project Glasswing, it was around 60% 150% more expensive than the current API costs.

From that link: "Afterward, Claude Mythos Preview will be available to participants at $25/$125 per million input/output tokens", compared to the current API price at $10/$50 per million input/output token.

u/IReportLuddites (56 points · permalink)

Beginning July 20th, most of you will continue using 5.6 Sol.

u/actkms (55 points · permalink)

And I still can't use it in health line of work because it flags every. single. prompt. It's ridiculous

u/Clueless_Nooblet (51 points · permalink)

Why even use it anymore at this point, now that there are two alternatives: Sol and K3.


2023 vs 2026

279 points · 60 comments · r/OpenAI · by u/EchoOfOppenheimer

Meme comparing 2023 and 2026 AI attitudes

A meme post comparing the AI discourse of 2023 with 2026, highlighting how the conversation has shifted from skepticism about AI capabilities to concerns about AI's societal impact and the normalization of AI tools in daily life.

Top Comments

u/matteoianni (85 points · permalink)

Where are all my "stochastic parrot" frens at?!!!

u/Tilting_Gambit (34 points · permalink)

They're out there. Some guy in another thread was calling AI a "plagiarism and CSAM" generator yesterday. There's just a lot of Luddites out there.

u/xikxp1 (16 points · permalink)

Well LLMs still struggle with 4th grade word problems. The ones that didn't appear that much in RL corpus

u/Frosty-Meeting-1606 (1 points · permalink)

Imo a lot of business will be wiped out assuming some will have access to frontier AI while others will not. AI boosts not only coding, but all kind of research and analysis tenfold.

If AI becomes expensive, we will see large corps annihilating smaller businesses much much quicker than the normal competitive pace

u/impatiens-capensis (4 points · permalink)

I think there's still a stochastic parrot argument to make. Math and code are just domains where examples are very easy to specify, generate and verify. So your stochastic model ends up generalizing further (but not completely) because it simply has so much data to work with and you can throw it at a bunch of problems and report on the ones it catches.


46 more Reddit stories

Updates: 05:30 AM PDT · 08:30 AM PDT · 11:30 AM PDT · 02:30 PM PDT · 05:30 PM PDT