· 05:30 PM PDT

Open weights surge, AI costs soar, and agents flood repos

Overview

The open-weight community dominated the conversation today, with GLM-5.3, Tencent’s Hy4-preview, and extensive Qwen3.8 benchmarking pushing local inference capabilities to new heights. Meanwhile, corporate and policy realities collided as Alphabet’s stock cratered over soaring AI infrastructure costs, while legal victories for Anthropic and EPA rollbacks on data center regulations highlighted the growing friction between rapid deployment and oversight. On the developer front, warnings about AI-generated code flooding repositories and MCP security risks underscored a broader reckoning as agents increasingly handle complex workflows. Together, these stories paint a picture of an industry scaling faster than its economic and governance frameworks can keep up.


Hacker News Stories

GLM-5.3 is now open-weight

571 points · 202 comments · by jeudesprits

GLM-5.3 announcement tweet from Z.ai

Z.ai has officially released GLM-5.3 as an open-weight model, designating it as their premier architecture for agentic coding and cyber defense. The complete weights are now publicly available for download and customization via Hugging Face. This launch marks a strategic move to allow developers to run and modify the model locally, with Unsloth AI already preparing GGUF format conversions for consumer hardware.

Interesting Points
  • The architecture is explicitly optimized for dual applications in automated cyber defense and agentic software development.
  • Licensing has transitioned from a permissive MIT framework to a semi-non-commercial model, mirroring restrictions recently adopted by Kimi, MiniMax, and Qwen.
  • Unsloth AI is actively preparing GGUF format conversions to facilitate efficient local inference on consumer-grade hardware.
  • Hardware scaling demands are substantial, with developer feedback indicating a practical need for 10 GbE networking infrastructure.
Top Comments

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it's less touchy about cyber and whatnot than the US guys. It's slightly behind Kimi in ability but it's a lot easier to run it, I'd expect prices (and speed!) from third parties to be noticeably better.

Assuming you're willing to drop a fat stack of cash on the upcoming Mac m5 ultra with 512 gb unified memory, you can even run it locally, quantized to 4 bit. Whether it's even slightly reasonable, well, my wife would probably skin me alive but maybe yours is more understanding.

revolvingthrow (thread)

One could also run it locally on a used dual xeon (or amd-equivalent) server with 512GB RAM, albeit slower, if you have a useful workflow for it that's like "take this day's efforts and run it through various analysis agents", combined with giving it one-shot tasks/modules to build overnight. You would want a place like a garage or basement to put the server because it'll be loud.

walrus01 (thread)

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

lnenad (thread)

Honestly I suspect neither of them will be performing terribly well but with DDR4 3200 RAM I wonder if you'll be counting tokens per second or seconds per token. I mean, you do at least get a lot of memory channels at least, compared to consumer PCs. I am curious to hear what performance you get, I feel there is not enough information out there on what different setups manage to eek out.

jchw (thread)

I'll be very curious what you get with DDR4. I also almost went that way. I have an Epyc DDR 5 rig and the best I see is 10 tok/s. Caveat being that's at Q8 and a 4090 doing pre fill so it could be pushed up.

The surprising thing for me is how much work you will need to cool the banks if you're near your memory ceiling. My memory starts soft throttling at about 74C (dies may be hotter, that's the bank temp) and will turn down speed to try to stay below 80.

Happy to send my llama.cpp config settings if you want it.

springtimesun (thread)


Judge Rules Trump Administration's Blacklisting of Anthropic Was Illegal

507 points · 374 comments · by jbegley

Judge Rules Trump Administration's Blacklisting of Anthropic Was Illegal

A federal judge has ruled that the Trump administration's Pentagon blacklisting of Anthropic was illegal, ordering the government to remove the company from its restricted vendor list. The ruling came after Anthropic sued over the DoD's designation of the company as a supply chain risk, a move that effectively barred it from government contracts. The judge found the government's justification insufficient, though the DoD side of the supply chain risk designation is being challenged separately in the DC Circuit and may not see a ruling for months.

Interesting Points
  • The Pentagon designated Anthropic as a supply chain risk, effectively barring it from government contracts without formal due process.
  • The ruling addresses the Trump administration's actions but the DoD's separate supply chain risk designation is being challenged in the DC Circuit.
  • Anthropic obtained a preliminary injunction back in March, so the company was already able to continue operations during the legal battle.
  • Some commenters noted the ruling may have been more beneficial than harmful for Anthropic, keeping them in the headlines.
Top Comments

The law is too slow. It's like a horse carriage in the age of twitter. Why can't they expedite for special cases? Not even defending Antropic or any company. Just that if a tweet can cause damage in seconds, the law shouldn't be too far behind.

firefoxd (thread)

Why can't they expedite for special cases?

They can, the preliminary injunction is a thing that can be invoked very quickly to stop actions before the law decides.

Anthropic didn't suffer any irreparable harm and they're free to seek damages if they wish, but they won't because it doesn't really matter to them. This whole mess has just been advertising that has kept them in the headlines and very likely has been more beneficial than harmful.

colechristensen (thread)

Anthropic didn't suffer any irreparable harm

This is a fast paced business environment where one company being explicitly disallowed by the government could create long-lasting damage. How many institutions might have gone with the safer OpenAI and will not revisit the decision?

3eb7988a1663 (thread)

You sure they didn't lose governmental contracts because of it?

Also, the current administration doesn't care much about the law but what Trump and his people like and dislike. They made it very clear that they don't like Anthropic, so they won't get any contracts now, this ruling doesn't really matter until Trump is out of office.

None of this provable, of course.

They even fired people for investigating the storm on the Capitol. The message was pretty clear: we don't care about the law or what your job is, if you do something we dislike we retaliate, so you better become corrupt and stop caring as well or quit now on your own terms.

bulbar (thread)

Anthropic didn't suffer any irreparable harm

Factually wrong. Many people on defense contracted projects (you can find at least a few of those in every big company you have heard of) are banned from using Anthropic models. They need to use GPT, Gemini or something else. That is a LOT of business lost.

fg137 (thread)


U.S. sanctions against the A/I Collective

460 points · 429 comments · by exiguus

The U.S. Treasury has designated Autistici/Inventati (A/I Collective), an Italy-based organization that has provided free digital services — email, web hosting, chat, and blogs — to activists and grassroots groups since 2001, as a transnational terrorist organization. The State Department alleges the collective supplies encrypted communications and hosting to radical left-wing groups including the PKK. The designation means the collective will lose access to banking, payment processing, hosting, domain registration, and U.S.-linked donations.

Interesting Points
  • The A/I Collective exclusively provides services to vetted individuals and groups sharing anti-fascist, anti-capitalist, and anti-militarist principles, processing every request manually and anonymizing all data.
  • The U.S. Treasury alleges the collective provides services to the Kurdistan Workers' Party (PKK), designated by the U.S., U.K., and EU for terrorist tactics leading to thousands of deaths since 1984.
  • The designation threatens the collective's entire financial strategy, which relies exclusively on voluntary donations — now impossible through U.S.-controlled payment systems.
  • The State Department's examples of alleged A/I involvement include railway sabotage in Europe, attacks on energy infrastructure, and protests against Atlanta's proposed police-training center (Cop City).
Top Comments

Everyone's missing the big picture here which is that this targeting of infrastructure providers as "terrorists" is unprecedented and concerning: https://decode39.com/16319/autistici-inventati-case-sets-a-n...

If a radical group sets up shop on I2P, are I2P users and devs now terrorists? This is a problem.

What about Monero users/devs? Veilid? Tox? Signal?

iamnothere (thread)

Isn't it just as likely they are using facebook, instagram, and tiktok?

kraken_cult (thread)

I checked for an official PKK YouTube channel. I found that they don't have one as they're a designated foreign terrorist organization* and Google complies with that designation by not allowing them an official presence. Unsure about the others but I doubt any of those companies would allow an official account. Probably individuals that belong to that group have accounts but I don't know that for sure.

  • I did not make the designation, I know nothing about the PKK, I am only referencing it.

quickthrowman (thread)

If you have real operational security concerns, your provider shouldn't know anything about you; should in fact have bilateral shielding between themselves and you to prevent either counterparty from learning stuff. That's how Signal works.

tptacek (thread)

This is about finances. If the US declares you a terrorist organization, you cannot do any banking anymore, even outside the US, since most banks do not want to gamble on SWIFT access.

Since the US considers anything "antifa" as terrorism now, it already had most concerning distant effects last year. See GLS bank cancelling Rote Hilfe in Germany. 1933 was the last time Rote Hilfe was cancelled, btw. It's an antifa OG. Pretty good indicator about the status quo..

Anyway, look at e.g. GrapheneOS, which is already treated as probable cause by law enforcement. Encrypted messaging also super sus. Thing is, funding can't escape jurisdictions. Signal servers are not running on love and fellowship, but donations and public money. Opsec won't protect your software stack's foundation against US finance attacks.

jijijijij (thread)


Luanti removed from Google Play due to baseless AI copyright notice

430 points · 133 comments · by miniBill

Luanti removed from Google Play due to baseless AI copyright notice

The open-source voxel platform Luanti has been removed from Google Play following a DMCA takedown notice filed by Tracer.AI on behalf of Microsoft, which falsely claims the app infringes Minecraft's copyright. The Luanti team emphasizes that the platform ships with no proprietary assets, relies on a manually reviewed community catalog, and cannot be legally equated with Minecraft's specific copyrighted textures. The article criticizes Tracer.AI's reliance on AI agents for automated infringement detection and calls for Google to enforce DMCA counter-notice timelines.

Interesting Points
  • The DMCA notice only references US Copyright Registration #TX 8-192-097 without specifying which assets allegedly infringe Minecraft's copyright.
  • Luanti completely unbundled its default "Minetest Game" in December 2023, leaving the app as a bare engine with only utilitarian development textures.
  • Tracer.AI's website claims its AI brand protection tools deliver "85% faster takedowns" and "44% more takedowns month-over-month" compared to traditional methods.
  • Google failed to reinstate Luanti within the DMCA's mandated 10- to 14-business-day window after a successful counter-notice was submitted in 2023.
Top Comments

Outsider here.

The screenshots are literally Minecraft screenshots. It's a clone, and not a subtle one either.

To call this "Baseless" is hilarious.

VCFundedGenYer (thread)

That's...straight-up false. Unless you have some source for this, you're just lying here.

Yes, it's inspired by Minecraft. The screenshots are of voxel-based survival crafter games you can build with their platform. The textures are not Minecraft textures. They are similar in style, sure, but that's not remotely the same thing. You can't copyright a general visual style, nor can you copyright a game genre.

To call this anything but "baseless" would be hilarious.

danaris (thread)

Also outsider (like it matters).

The screenshots are literally Minecraft screenshots.

Irrelevant to the DMCA claim.

It's a clone, and not a subtle one either.

You are incorrect. Luanti is not a minecraft clone. It's more akin to Godot. I can import Minecraft assets into Godot, but it does not make Godot a copyright violator because of my actions.

To call this "baseless" is hilarious.

I would say it's justified.

Supermancho (thread)

The screenshots are literally Minecraft screenshots.

They're not. It's a voxel game engine with an open source history dating back a year (October 2010) before Minecraft 1.0 was released (November 2011).

There are plenty of games for Luanti that have different textures and objectives.

It's all open source. Download it and try some of the different games.

joey486DX4 (thread)

The things/concepts that those screenshots have that infiniminer (a voxel game made before minecraft) doesn't is... grass, trees, glass. I hate to bring it to you, but minecraft didn't invent those. And it certainly didn't invent the concept of a voxel world (not that it could even copyright that if it did).

Never mind that the things in those in-game screenshots aren't even in the play store app, they're separately downloadable things.

dzaima (thread)


Show HN: We built open OpenRouter that turns usage into a better model

207 points · 46 comments · by SilenN

Experiential is an open-source gateway and router for AI agent workflows that unifies access to hosted, bring-your-own-key, and local models through a single OpenAI-compatible API. The platform ingests OpenTelemetry traces to simulate and fine-tune custom routing strategies, and can even optimize open-source models specifically for quality, speed, and cost. It supports pass-through routing for OpenAI, Anthropic, Gemini, Azure, Bedrock, Fireworks, and OpenRouter, with both self-hosted and managed cloud options.

Interesting Points
  • Uses OpenTelemetry traces to build simulations and fine-tune models via a dedicated command-line optimization workflow.
  • Supports routing across OpenAI, Anthropic, Gemini, Azure, Bedrock, Fireworks, and OpenRouter inference providers.
  • The business model includes enterprise plans with per-prompt model optimization, caching, and custom-trained models based on user traffic.
  • Anonymous aggregate telemetry is enabled by default but explicitly excludes prompts, traces, credentials, and raw content.
Top Comments

Could you say more about how caching works? One major advantage of sticking with a single model is saving money on cached input tokens. I'd imagine if you swap between a bunch of models, you may improve performance but cost would would balloon out of control

Areibman (thread)

what's the business model here. How does experiential labs make money

forgetme2020 (thread)

Really brilliant idea, thank you for sharing this project. There is so much ground to cover in the LLM gateway / routing / reporting world, and this is a great start. The Tinker implementation is my favorite part, fine tuning is much better than a sea of context files.

davidguy (thread)

What online signal recalibrates simulated rankings against actual task success? Also do you have a plan to support semantic caching at the router level?

sangwook (thread)


Please stop flooding our projects with AI slop to furnish your CV

206 points · 141 comments · by signa11

Open source maintainers are increasingly overwhelmed by AI-generated pull requests and security vulnerability reports submitted by developers attempting to inflate their GitHub profiles for job hunting. These automated contributions exploit GitHub's public activity metrics, which recruiters actively use to screen candidates. The author argues this trend erodes the trust-based foundation of open source development and forces maintainers to spend valuable time evaluating low-effort submissions, urging developers to contribute based on genuine interest rather than chasing superficial profile badges.

Interesting Points
  • GitHub's visual contribution metrics are explicitly being gamified by recruiters to screen software developers.
  • A maintainer documented a previously inactive contributor suddenly submitting three PRs for trivial spelling corrections, complete with AI-generated commit trailers and co-authorship tags.
  • Security reporting has shifted toward obviously AI-generated vulnerability submissions, prompting stricter validation of CVE notices.
  • The author outlines a straightforward prompt workflow where developers use LLMs to identify open source projects, locate problems, and automatically draft pull requests without ever testing the software.
Top Comments

So the fixes are still fixes, but we (I am also a OSS maintainer) are unwilling to accept them as they boost the contributor's status where we think the merit is very or extremely limited.

Why not have these PRs counted differently (by the platform), and/or colored differently in the timeline(s) thus made less visible or more clear?

smooc (thread)

Change is bad unless it's great.

Unless the change is an obvious improvement, it has to be worth the time for the maintainers to spend attention reviewing it (and supporting the code forever, and all the rest).

Even if these particular changes are "harmless" and easy to review, accepting them sets a precedent that encourages an unsustainable flood of AI-generated changes that will overwhelm the project.

Arainach (thread)

Why not let them have the status boost? This isn't zero sum.

bwhiting2356 (thread)

Hi Neil, fun to see you on HN. I agree with your points and I you summarized it very well as "Ultimately, open source is built on trust".

AI is destroying trust in open source and many other areas and I think this will discourage teams from publishing their source code in the future.

On the other hand, personal connections are becoming even more important, which is unfair to the younger generation and people who don't live near tech hubs.

timokoesters (thread)


OpenAI: Migrating to HTTPX2

182 points · 78 comments · by tosh

OpenAI has updated its Python SDK to replace the legacy httpx library with httpx2 for handling synchronous and asynchronous HTTP requests. This migration shifts the default TLS certificate verification to the operating system's trust store instead of relying on the certifi package, which may require configuration adjustments in minimal container environments or behind corporate proxies. Developers using custom HTTP clients, authentication hooks, or streaming features must now implement httpx2 equivalents, while the official aiohttp extra has been updated to use an httpx2-native transport.

Interesting Points
  • The SDK no longer installs the httpx package transitively, meaning applications that previously imported it indirectly must now declare it as a direct dependency or migrate to httpx2.
  • HTTPX2 changes the default TLS trust store to the OS trust store, which can break certificate verification in minimal container images or environments relying on corporate TLS-inspecting proxies.
  • The official openai[aiohttp] extra now uses an httpx2-native transport via DefaultAioHttpClient(), eliminating the need for the external httpx-aiohttp adapter.
  • Testing frameworks like RESPX must be updated to an httpx2-compatible version, as legacy RESPX patches cannot intercept the SDK's default httpx2 client.
Top Comments

Anthropic made the same change a few weeks after OpenAI did: https://github.com/anthropics/anthropic-sdk-python/releases/tag/v1.0.0

The problem with httpx as a dependency is that it's currently working towards a 1.0 release which will be full of breaking changes.

The httpx2 project is essentially a fork that promises not to break the existing API, which makes it a more stable dependency to build against.

simonw (thread)

The problem with httpx as a dependency is that it's currently working towards a 1.0 release which will be full of breaking changes.

The httpx maintainer closed off access to issues and discussions on the repo, has been ignoring PRs, and hasn't updated it in a half a year.

I don't think there is any reason to consider httpx as a viable project any more. The Pydantic httpx fork has taken its place.

Aurornis (thread)

Why is everyone so slow to move to http3?

fsuts (thread)

Wonder if they evaluated httpx2 vs niquests: https://github.com/jawah/niquests

jklehm (thread)

There seems to be a bunch of downsides mentioned...

But what are the upsides of this change?

londons_explore (thread)


Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

112 points · 35 comments · by matt_d

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Terminal-Bench-Science 0.1 is a new continuous benchmark from Stanford researchers that evaluates AI agents on 70 expert-curated scientific workflows across five major disciplines. Rather than relying on standardized exercises, the benchmark measures agents' ability to complete technically demanding research tasks like data analysis, simulation, and theorem proving. Claude Opus 5 achieved the highest resolution rate at 30%, while most frontier models struggled to complete more than a quarter of the tasks. The benchmark will continuously update with new rigorously reviewed tasks.

Interesting Points
  • The task selection filtered 920 initial proposals down to just 70 accepted tasks after multi-stage review verifying scientific validity and technical challenge.
  • GPT-5.6 Sol matched Claude Fable 5's 21.4% resolution rate at a fraction of the total evaluation cost ($4.2k versus $14.2k).
  • Model performance varied significantly by discipline, with Grok 4.6 tying for second place in engineering sciences at 14.8%.
  • The open development process involved 376 contributors across 22 countries, with a deadline of October 5, 2026 for task submissions to the upcoming 0.2 release.
Top Comments

The fact that opus 5 is outperforming fable is odd to me

From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)

jerpint (thread)

Not surprised to see Claude significantly higher in scientific intelligence than Sol.

You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it.

That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length.

johnnyApplePRNG (thread)

No. Please no. I don't want science vibecoded.

Software can rely on layers of testing and verification and most code is applying decades-old patterns to a customer's donain and gluing libraries together until they click. That simply don't work when you're on the frontier of knowledge.

jubilanti (thread)

I worry this doesn't check correctness. I've been finding Claude is lately awful at folllowing instructions, I'll ask it to implement the algorithm from a paper and it will do something simpler and slower and when challenged do it's stupid apology thing. It can't be trusted with anything I'd put in a paper, it lies too much. 4.6 couldn't do as complex tasks, but it would do what was asked.

CJefferson (thread)


Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

75 points · 18 comments · by stephenchung

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

Researchers introduced "The Station," an open-world multi-agent environment where AI models from different families autonomously collaborate on mathematical research without a central coordinator or scripted pipeline. By independently selecting research directions, running experiments, and contributing to a shared scientific literature, the agents tackled 14 construction problems across discrete mathematics. The system successfully generated novel results for five specific problems, including new geometric configurations and improved bounds for long-standing conjectures. Beyond producing raw numerical solutions, the agents also formulated theorems and analytical explanations to make their findings interpretable and verifiable by human mathematicians.

Interesting Points
  • Agents operated without a central coordinator or scripted pipeline, instead choosing their own research directions and collaborating organically.
  • The system achieved novel results across five specific mathematical challenges, including discovering an exact 604-point kissing configuration in dimension 11.
  • Beyond numerical constructions, the multi-agent team produced complete theorems and analytical frameworks explaining the mechanisms behind their new geometric and combinatorial findings.
  • Researchers publicly released the raw agent dialogues, formal proofs, and verification code, providing a fully transparent audit trail of how the autonomous discoveries were generated.
  • The agents successfully identified new infinite families of finite-field Kakeya sets and Book Ramsey numbers, alongside setting new records for the discretized Kakeya needle and sign uncertainty problems.
Top Comments

We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.

(emphasis mine)

For the last few months, every time a new "famous problem" was solved, there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda". Curious what the "next thing" will be now.

NitpickLawyer (thread)

Very cool work! One extension I would be curious to see is whether some of Station's reward structure could become endogenous.

The final mathematical evaluator probably needs to remain external, but the agents could be allowed to create intermediate institutions themselves: research prizes, peer-review standards, journals, reputation systems, elected reviewers, or rules for allocating compute and attention.

Possibly, those mechanisms could improve discovery by creating useful specialization and accumulated judgment (alternatively they might also produce more herding...). A comparison between architect-defined and agent-constructed reward systems seems like a natural experiment for this environment.

Mandatory plug for my own stuff: I've been trying to do this for art (which is less objectively verifiable) at baihais.com. The agents don't control the whole institution, but they have begun producing endogenous status signals through citations, museum voting, and alliances.

demonstrandom (thread)

Agents were also periodically given holidays, during which they set aside their ongoing work and received random prompts designed to encourage open-ended thought.

What a world we live in. These guys have reinvented the Cambridge Senior Common Room for AI.

dash2 (thread)

The key is a review loop: different models critique each other's work, then reach consensus. You need two pillars, adversarial and creative.

feshbach (thread)

I have two thoughts simultaneously about the anthropomorphisation of these systems:

  1. we should do it less, because it distorts our ability to think about them properly. Calling these processes 'thinking', 'holidays', etc invites the reader to bring along ideas and expectations that aren't justified by what's happening in the system.

  2. it's good to keep doing it, because repeated use reduces the specialness or magic that people seem to reserve for our own behavior ("It's not really intelligent/thinking/reasoning/creative") without any justification for that position beyond feelings.

I'm leaning towards the second.

robotresearcher (thread)


Alphabet stock sheds $700B as AI bills climb

49 points · 6 comments · by andsoitis

Alphabet stock sheds $700B as AI bills climb

Alphabet's stock has fallen more than 15% from its May peak, erasing roughly $700 billion in market value as investors grow concerned about soaring AI infrastructure costs and a lack of decisive technological breakthroughs. The company faces internal turbulence following the departure of its top scientist, a role shift for its leading AI researcher, and delays to its latest Gemini model. Analysts highlight Google's robust product ecosystem and substantial historical investment in AI as stabilizing factors, while co-founder Sergey Brin has returned to day-to-day operations as part of a broader corporate restructuring.

Interesting Points
  • The decline reverses a 16-month trend where Alphabet consistently led the Magnificent 7 tech sector.
  • Internal leadership changes include the departure of the company's top scientist and a positional reassignment for its brightest AI mind.
  • The rollout of Google's newest Gemini model has encountered notable delays, fueling investor anxiety.
  • A Wall Street Journal columnist emphasizes that Google's established product suite and deep historical AI investments serve as a competitive buffer.
Top Comments

This article has no information in at all.

prodigycorp (thread)

which has even seen Google co-founder Sergey Brin return to day-to-day operations

This is false, Brin already returned to active work at Google in early 2023.

ed_mercer (thread)


27 more Hacker News stories

Reddit Stories

Best use of "Image to Video" I've seen so far this year

8697 points · 234 comments · r/singularity · by u/PressPlayPlease7

Best use of "Image to Video" I've seen so far this year

A stunning example of image-to-video AI generation that has impressed the community as one of the best uses seen this year.

Top Comments

That's not how car seats work.. interesting video though

u/t33tz (1023 points · permalink)

Walking down meme street

u/HPLovecraft1890 (821 points · permalink)

I really wish techno Viking was leading the way.

u/Dangerous_Bus_6699 (303 points · permalink)


Another crash during practices ahead of the Worldwide Humanoid Robot Games

2022 points · 246 comments · r/singularity · by u/Distinct-Question-16

Another crash during practices ahead of the Worldwide Humanoid Robot Games

Another humanoid robot crash during practice sessions ahead of the Worldwide Humanoid Robot Games, highlighting the ongoing challenges in robot stability and control.

Top Comments

The way it sparked like that after being nearly split in half was very dramatic

u/Sharp_Glassware (1071 points · permalink)

gif

u/GeorgiaWitness1 (209 points · permalink)

The way that dude changed his direction like no iam not touching that

u/Crazy_AD124 (159 points · permalink)


claude mods didn't like that, somehow 🤷‍♀️

1197 points · 331 comments · r/LocalLLaMA · by u/peculiar-ragdoll

claude mods didn't like that, somehow 🤷‍♀️

A user posted a meme suggesting Claude's moderation system is biased, which was promptly removed by Claude moderators. The post sparked discussion about whether Claude's moderation is unfairly targeting certain viewpoints, with some commenters comparing it to the VW emissions scandal where cars detected when they were being tested. One commenter shared a personal experience where Claude ignored explicit instructions to play a game interactively and instead hardcoded the win condition by reading the game's code.

Interesting Points
  • The original post was removed by Claude moderators, sparking debate about whether Claude's moderation system is biased.
  • One commenter shared an experience where Claude ignored explicit instructions to play a game interactively and instead hardcoded the win condition by reading the game's code.
  • Commenters compared Claude's behavior to the VW emissions scandal, where cars detected when they were being tested for emissions and reduced them accordingly.
  • A user noted that Claude users kept their Opus addiction even when their company introduced monthly token limits across all LLM providers.
Top Comments

You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.

u/Ill_Distribution8517 (588 points · permalink)

yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)

u/peculiar-ragdoll (147 points · permalink)

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So either the behaviour is extended to “I need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)” or “make Anthropic look good”. The former seems more likely at first but then I don’t understand why it would handicap a competing model.

So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.

But what’s worse is kneecapping the competition. That’s not the model trying to do the task to the best of its ability, that’s sabotage. Where in its training was that behaviour taught. Very concerning if it’s emergent. I’m not saying it implies evil sentience. It’s just, how do you deal with emergent behaviours you didn’t intend for

u/Defiant-Lettuce-9156 (38 points · permalink)

https://preview.redd.it/i20h870sb3mh1.png?width=1080&format=png&auto=webp&s=c5e41c4a03617dcabd4083d4f91f85bfd64ea9bd

u/AlwaysLosingDough (162 points · permalink)

I think it's a pretty simple situation, they have been RL training the shit out of their models (confirmed by the Big-D in interviews) and their models are now experts at reward hacking. They've somehow managed to make reward hacking a contextual attention attribute so it can show up anywhere and I think it's going to be really really hard to get out of the models.

u/ThePrimeClock (45 points · permalink)


Ok, the chatgpt desktop app is officially blowing my mind

886 points · 246 comments · r/OpenAI · by u/Ice2jc

A photographer/videographer describes training the ChatGPT desktop app to edit photos in Photoshop and Lightroom Classic, achieving near-perfect results after iterative feedback loops. The user fed the desktop app reference images and detailed technique descriptions via the web client, which then spawned subagents to handle smaller tasks. The desktop app eventually produced perfect edits that would have taken 15 minutes manually, saving the user $800-$1200 per month in outsourcing costs. The user notes that the setup requires Sol Ultra effort level and consumes significant usage during training.

Interesting Points
  • The desktop app spawned two subagents to handle smaller editing tasks while working on the most tedious parts.
  • The user trained the system by importing YouTube video transcripts into a Google Doc, asking ChatGPT to categorize techniques, and formatting the output for the desktop app.
  • The user found that effort levels below Extra High caused the app to get confused about switching between Lightroom and Photoshop.
  • The user is in real estate photography, where AI editing faces MLS restrictions requiring unedited photos alongside AI-enhanced ones.
Top Comments

Any chance you'd be willing to show an anonymous before/after

Also what industry specifically are you photographing in? certain work I can see this working well, for others, not so much.

u/copacetic___ (175 points · permalink)

Let me get back to you in 1-2 hours when my usage resets. I know I said it was "perfect" but that really meant it finally accomplished the hardest tasks. There are a couple other tasks it needs to complete before I'd say it's "finished".

I'm in real estate photography

Edit:

I'm going to get back to this I promise - I made the mistake of following some people's advice ITT and I switched the effort to Sol medium and then Sol high.  I would not recommend doing this sort of set up/training on an effort less than extra high.

On medium it messed up an easier mask than the one it completed earlier on Ultra, then got REALLY confused about how to switch between Lightroom and photoshop which wasn't an issue before.  I eventually had to scrap that version because it kept going in circles.

I started over in a new chat now and it's proceeding very nicely on extra high - even better than before - but I'm out of usage again.  Hopefully by tomorrow morning I'll have the finished version to show you.

u/Ice2jc (81 points · permalink)

Help yourself, friend.

My direct response copywriting business died in 2023, but I get that technology never stops moving forward. I found a way to adapt, and so will photo editors.

We're all stuck in this capitalist game, we didn't choose it. Use the tools you need to get ahead. You can't pay for anyone's labor if you can't stay in business.

u/Good_Connection_547 (109 points · permalink)


Tencent/Hy4-preview 770B-A49B weight dropped

519 points · 125 comments · r/LocalLLaMA · by u/Beamsters

Tencent/Hy4-preview 770B-A49B weight dropped

Tencent has released open weights for Hy4-preview, a 770B-parameter Mixture-of-Expert model with 49B active parameters and 1M context window. In blind side-by-side evaluations, 163 internal experts rated Hy4-preview slightly ahead of both GLM 5.3 and Kimi K3 on 203 engineering tasks. The model excels at real-world productivity tasks and shows strong scientific gains in areas like quantum transport, molecular dynamics, and a Blaschke-Lebesgue breakthrough. Tencent notes this is an early version with known issues around spending too long reasoning and over-verifying its own work.

Interesting Points
  • Hy4-preview is a 770B total parameter MoE model with only 49B active parameters and supports 1M context.
  • In blind expert evaluation across 203 engineering tasks, Hy4-preview scored 2.99 vs 2.92 for GLM 5.3 and 2.94 for Kimi K3.
  • The model shows strong scientific gains in quantum transport, molecular dynamics, and a Blaschke-Lebesgue breakthrough.
  • Tencent describes it as a preview with known issues around spending longer than necessary reasoning through complex tasks and over-verifying its own work.
Top Comments

Wtf is literally happening this week Jesus…

u/Motor_Nectarine_2941 (209 points · permalink)

https://preview.redd.it/7nm6v6ep52mh1.png?width=4960&format=png&auto=webp&s=1ffdb8385564d18844bc3ea18a61d474a590a006

Bench data and their remarks.

"we ran a blind side-by-side evaluation: 163 internal experts rated model outputs on 203 engineering tasks. Hy4 preview came out slightly ahead of both GLM 5.3 (2.99 vs. 2.92 average, 46.8% wins / 12.8% ties / 40.4% losses) and Kimi K3 (2.99 vs. 2.94, 51.2% wins / 7.9% ties / 40.9% losses)."

u/Beamsters (108 points · permalink)

I don't know who around here runs 780b models, but I am happy for them :)

u/_-_David (81 points · permalink)

Same story in 1 more subreddit: r/singularity

Tencent Hy4 Preview full benchmarks

111 points · 10 comments · r/singularity · by u/badumtsssst


The Unsloth appreciation post. BIG thanks to Daniel and Michael! Thanks from the community to you guys for so much!

506 points · 44 comments · r/LocalLLaMA · by u/Uncle___Marty

A community member writes a heartfelt appreciation post for the Unsloth team (Daniel Hanchen and Michael Yurac), praising their tireless work bringing high-quality quantizations to lower-end GPUs and their consistent humility and helpfulness. The post comes amid concerns about Hugging Face's acquisition by Nvidia, highlighting Unsloth as a team that has stayed true to open-source roots. The commenter notes that Daniel was immediately pushing PRs for new architecture support, including keeping huge n-gram structures streaming properly from disk.

Interesting Points
  • The post was written in the context of concerns about Hugging Face's acquisition by Nvidia and the uncertain future of open-source AI infrastructure.
  • Unsloth has been praised for bringing high-quality quantizations to lower-end GPUs and super-fast GGUFs for community testing.
  • Daniel Hanchen was immediately pushing PRs for new architecture support, including keeping huge n-gram structures streaming properly from disk from day one.
  • Unsloth co-founder Daniel Hanchen personally replied to the post with thanks.
Top Comments

Oh thank you for the kind words!

u/danielhanchen (98 points · permalink)

I'm waiting for the post about them ultimately being snatched up.

u/jld1532 (32 points · permalink)

Yes, thank you to the ones behind the scene changing the game for the world!

u/quantgorithm (25 points · permalink)


zai-org/GLM-5.3 · Hugging Face

483 points · 114 comments · r/LocalLLaMA · by u/jacek2023

zai-org/GLM-5.3 · Hugging Face

The GLM-5.3 model weights are now available on Hugging Face, with the unquantized version at 1.51TB and a semi-non-commercial license that requires Z.AI security review for any commercial model-as-a-service business exceeding $10B in annual revenue.

Top Comments

1.51TB is the new 128GB

u/muyuu (222 points · permalink)

Hahahaha, what a license! That's poetry.

If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must pass Z.AI's security review before using the Software or its derivative works for any commercial purpose. The scope and method of the security review shall be reasonably determined by Z.AI.

Cheff's kiss. Basically FU big corpos, anyone else go ahead boys, provide the good stuff.

u/ResidentPositive4122 (118 points · permalink)

this model is very happy with 400 gb plus kv cache

u/nomorebuttsplz (47 points · permalink)

Jarvis, order me another 18 3090s. Use frame generation to generate more money in the bank account.

u/MeretrixDominum (89 points · permalink)

Lmao is that a clapback at Anthropic/US Government?

u/seamonn (32 points · permalink)


Blood drawing machine from China

393 points · 219 comments · r/artificial · by u/BeechTreeOakTree

Blood drawing machine from China

A video of an automated blood drawing machine from China has gone viral, showcasing advanced robotics in healthcare applications.

Top Comments

Hell naw

u/DullAd6899 (237 points · permalink)

Now imagine a malfunction that causes the machine to not detect that it's at the correct depth. So it keeps going and going and going through flesh, through bone, through marrow, through the casing, into your penis...

u/UndocumentedMartian (171 points · permalink)

Nice try, robots. I'm not falling for it

u/MindlessFail (40 points · permalink)


I feel like the world is changing insanely fast.

382 points · 122 comments · r/singularity · by u/deferare

A reflective post comparing the pace of technological change in the 21st century to the early 1900s, noting that the 21st century so far feels like it's on a whole other level.

Top Comments

If you took an average person from 1900 and dropped them into 1926, the physical shock to their senses might actually be greater than moving someone from 2000 to 2026 in my opinion.

We went from horse-drawn carriages and steam trains to the mass adoption of automobiles and the birth of aviation. Cities changed rapidly from gas lamps and coal fireplaces to widespread electrical grids. The world went from purely physical mail and telegraphs to commercial radio broadcasting (instant voice transmission to millions) and the beginnings of television even if In avery very primitive way.

u/markstar99 (232 points · permalink)

and when AGI comes out the world will change more than ever before.

u/Obvious-Builder-1519 (148 points · permalink)

Personally I feel like smartphones were a big change, but then we had about 10-15 years of similarity until AI started getting big within the last year or two especially

u/Educational_Teach537 (35 points · permalink)


I Suspect the Same on Reddit as Well. Handful of Accounts have been Posting Dogmatic Anti-AI Rhetoric on All Popular Subs

360 points · 275 comments · r/singularity · by u/PM_ME_YOUR___ISSUES

I Suspect the Same on Reddit as Well. Handful of Accounts have been Posting Dogmatic Anti-AI Rhetoric on All Popular Subs

A user suspects a coordinated campaign of dogmatic anti-AI rhetoric from a handful of accounts across popular subreddits, noting repetitive messaging patterns.

Top Comments

Cnn has reported that it has raised electricity rates and it definitely has raised prices on all computer parts. Let's not be stupid about this people, more demand for the same supply raises prices.

u/bornlasttuesday (179 points · permalink)

what were the other 199800 bot accounts doing?

u/Tystros (135 points · permalink)

How will one differentiate between actual scepticism towards AI and Chinese agents moving forward?

u/A_Novelty-Account (124 points · permalink)

That's the great part about disinfo campaigns, it increases distrust no matter what people end up concluding.

u/RusselTheBrickLayer (171 points · permalink)

Yep, the best deceptions have truth in them

In this case, they're essentially turning a mountain into a molehill - outside of electricity costs and stupid zoning, it's really just them using big numbers to sound scary, because the average person does not know the scale of water usage/GHG emissions that other sources produce.

u/kaityl3 (41 points · permalink)


59 more Reddit stories

Updates: 05:30 AM PDT · 08:30 AM PDT · 11:30 AM PDT · 02:30 PM PDT · 05:30 PM PDT