Qwen 3.8 Dominates Local AI as Humanoid Robots Shatter Records
Overview
Qwen 3.8 27B has taken over developer forums, delivering frontier-level coding and reasoning capabilities on modest local hardware while sparking intense debate over model quantization and VRAM optimization. This surge in accessible AI is matched by rapid progress in robotics, with humanoid machines recently shattering human world records in sprinting and autonomous tennis. Meanwhile, industry executives are recalibrating expectations around AI adoption speed and enterprise data dependency, even as major labs adjust pricing and release new desktop tools. The day’s conversation underscores a clear pivot: highly capable AI is no longer confined to massive cloud clusters, and its physical applications are accelerating faster than anticipated.
Hacker News Stories
'AI refuser' quit her dream job, and hopes others follow
34 points · 39 comments · by mcapodici
Gabrielle Boyle, a former AFL participation manager, resigned after the league refused to exempt her from using Microsoft Copilot AI, highlighting a significant gap in Australian workplace law. Current employment and privacy regulations generally allow employers to mandate AI tools on company equipment without offering workers an opt-out. While unions argue the Fair Work Act requires meaningful consultation when AI changes job functions, legal experts stress that individual exemptions remain unenforceable and ineffective at preventing broader data collection. The case illustrates how poorly managed AI implementations are driving away skeptical staff and exposing urgent needs for updated governance frameworks.
Interesting Points
- An internal survey revealed 35% of AFL staff were uncomfortable with AI adoption, yet more than half of Australian Services Union clerical workers were unaware of their employer's AI policy entirely.
- The AFL's official AI policy explicitly endorses ChatGPT, Copilot, and Claude Code, while a footnote admits the policy document itself was drafted with generative AI assistance.
- New South Wales recently passed the Work Health and Safety Amendment (Digital Work Systems) Act, though its application to general-purpose assistants like Copilot has not been legally tested.
- Research by KPMG and the University of Melbourne shows Australians rank among the least trusting of AI globally, with skepticism driven primarily by employer conduct and change management.
Top Comments
I really hate how the excruciatingly overdue reckoning for tech elites and income inequality is turning into this religious schism over AI and datacenters.
We're going to smash the looms rather than figure out how to share them, again.
— cagenut (3 replies)
The underlying issue is really capital ownership, what classifies as capital, protections around it, and how much leverage capital provides on a society and democracy (that part being the most important, IMHO).
Who owns the looms and what they do with them isn’t inherently an issue if all loom owners can do is buy an extra Yacht. Instead they can enforce ungoverned law on society through a combination of disproportionate influence in government and through market forces where private policy (especially at large) become nearly undifferentiated from law (the policy that benefits them becomes so widespread and normalized that alternatives are for all intent and purposes impractical or unreasonable, therefor private policy within a capital ownership domain is law or they’re a monopoly so their policy is the policy).
But I think you’re right, as always we’re going to focus on the adjacent issues vs addressing the root of the problem. The issue is what wealth inequality affords one, when it is capable of infringing on rights and livelihoods of others, not that they necessarily have to share the other luxuries and rewards of their attained wealth. I care not how many luxuries in life Musk has, I may be envious from time to time but whatever. I care a lot more when things he does or says has unrealistic influence and affects me directly, just because he sits atop a mountain of capital and we pretend that mountain of capital somehow was bestowed upon him from divinity that he should have such influence. I’m picking on Musk because he’s the richest and has clear examples of this, he’s by no means alone… it’s that class of wealth at large.
— Frost1x (0 replies)
I don't appreciate the celebration around acts of individualism rather than collectivizing to improve our workplace for all.
— wildrhythms (3 replies)
AI boosters on this site will be very mad that this person exists, but I'm glad there are people that stick to their principles. Principles that aren't just "productivity" or "efficiency". People can and should have other motives that are just as if not more important.
— tkel (2 replies)
Does Australia have privacy provisions similar to the gdpr?
— CalRobert (2 replies)
Software Engineering in the Agentic Era
26 points · 9 comments · by silverpiranha
Simon Willison is launching "Agentic Engineering Patterns," a continuously updated guide documenting best practices for professional software engineers using autonomous coding agents. He defines "agentic engineering" as a discipline where engineers leverage tools that can both generate and execute code to test and iterate independently, contrasting it with "vibe coding." The project is structured as a series of evergreen chapters inspired by classic software design patterns, with the first two already published. Willison plans to release one to two new chapters weekly while maintaining a strict policy that all written content remains human-authored.
Interesting Points
- The first two published chapters examine how near-zero initial code generation costs disrupt traditional engineering intuitions and how red/green test-driven development helps agents write more succinct, reliable code with minimal prompting.
- Willison's existing archive of AI-assisted programming posts has already surpassed 345 entries, highlighting the need for a centralized, structured resource.
- The Django backend and views powering the new guide format were almost entirely written by Claude Opus 4.6 running in Claude Code via an iPhone.
- Unlike standard blog posts, the guide's chapters are designed as evergreen content that is updated over time rather than being frozen at first publication.
Top Comments
A few nits:
- Writing code is cheap now
Change to "generating code is cheap", reserve "writing" for the manually written code for the pre-AI era. I think this is a good wording separation.
- Agentic Engineering Patterns
I must ask to add at least a chapter to be read by agent. I.e., patterns just to tell agents how human might be working when working with them. Without this, I believe the book's content will be less relevant in 3 months, but with that, it feels a agentic-native book to me. (this is not try to be cute, we have to write for agents now)
— bigcat12345678 (0 replies)
one theory I am thinking through is that people say AI is good at greenfield and has a harder time editing code in a legacy code base.
So maybe you just treat projects as greenfield instead of editing.
Like if you want to make changes to a page you generate a new version of the page instead of editing the old one in place.
Then you keep the old version as fallback if problems come up with the new version.
— pianopatrick (1 reply)
Palantir's Karp – frontier AI labs that are 'trying to drug addict us'
19 points · 8 comments · by rishabhd
Palantir CEO Alex Karp warned that frontier AI labs are creating dependency among enterprises, arguing that companies need to protect their data or risk losing their business to model makers. He drew a parallel to Chinese models distilling U.S. models, noting that frontier labs have already distilled all the value of intellectual property everywhere. The comments suggest the argument, while potentially valid, is self-serving since Palantir's own forward deployed engineering and consultancy model faces the same threat from frontier models.
Interesting Points
- Karp argued that enterprises need to protect their data or risk losing their business to model makers.
- He said Chinese models can't be blamed for distilling U.S. models when the frontier labs "distilled all the value of IP, everywhere."
- The comments suggest the argument, while potentially valid, is self-serving since Palantir's own forward deployed engineering and consultancy model faces the same threat from frontier models.
Top Comments
*"Karp said enterprises need to protect their data or risk losing their business to model makers.
Karp said Chinese models can’t be blamed for distilling U.S. models when the frontier labs “distilled all the value of IP, everywhere.”*
I think all of us would echo that.
— andsoitis (2 replies)
While the argument makes sense, the motivation is rather self-serving: he is afraid frontier models will steal his forward deployed engineering/consultancy model.
Palantir is stealing business from consultancy firms, now frontier models are going to do the same to them. He has to convince executives (main purchase decision makers) by what he does best: politics.
— mgh2 (0 replies)
Sauron warns against Saruman.
— andrewstuart (0 replies)
Andrew Ng: AI Engineering Skills Map: Building and Deploying AI Applications
15 points · 0 comments · by Anon84
Andrew Ng expands on his AI Engineering Skills Map by detailing the core competencies required for building and deploying AI applications. He argues that the fundamental challenge of AI engineering lies in the inherent unpredictability of model outputs, which transforms development into a highly iterative, experimental process rather than a linear one. To navigate this uncertainty effectively, Ng outlines six essential technical domains: LLM foundations, data grounding, agentic system design, evaluation-driven development, production operations, and machine learning fundamentals.
Interesting Points
- The skills map was constructed by analyzing a large volume of job postings, conducting structured expert interviews, and reviewing survey responses.
- Understanding LLM internals like tokenization and context windows helps engineers decide when to rely on models, manage tradeoffs in knowledge cutoffs, and optimize for cache hits and sampling parameters.
- Beyond early vector search RAG, grounding techniques now include knowledge graphs and semantic layers over structured data, requiring engineers to decide what to hardcode in prompts versus what to retrieve on-demand via tools.
- Ng identifies driving a disciplined evals and error analysis loop as the single most distinguishing trait of top AI engineers, emphasizing the need to blend deterministic code-based checks, LLM-as-a-judge methods, and human-in-the-loop feedback.
- Operating AI in production demands statistical regression testing calibrated to risk, alongside continuous monitoring for model drift, adversarial prompt injections, and cost/latency optimization through techniques like model distillation.
Why can AI generate Super Mario but not a wedge ramp for my robot vacuum?
11 points · 5 comments · by zhuchaokn
A discussion on Hacker News about the surprising gap between AI's ability to generate recognizable cultural artifacts like Super Mario and its difficulty with practical engineering tasks like designing a wedge ramp for a robot vacuum. Commenters suggest that parametric design tools like OpenSCAD, combined with LLMs, can produce surprisingly good results for part design when given metric units and sketches as context.
Interesting Points
- One commenter reports getting Claude to generate a complex mating part from a 3D scan using OpenSCAD, with back-and-forth refinement
- Another commenter points out that the real challenge isn't modeling but decomposing problems into usable chunks and anticipating what will be a problem
- A developer shares Aetheris, an open-source geometric modeling kernel for LLM-based CAD, noting it competes with ecto's vcad and CadQuery/Build123d approaches
Top Comments
Ive had great success with asking claude to use openscad for part design.
Because it's parametric design as opposed to modeling, it inherently lends itself to making it easier to build.
Use metric units when prompting, and Ive even drawn sketches to add to context, and be surprised at how good the outcome can be!
— chews (thread)
modeling isn't the hard part. It's decomposing it into usable chunks, the knowledge of what's going to be a problem, etc.
either way, give an actual CAD program like fusion a go. No, not the FOSS stuff- the real stuff. Although, some of them are "close, but no cigar" these days. Not like gimp which is an exercise in self-flagellation.
— butvacuum (thread)
Well, because you need a geometric modeling kernel to generate BRep CAD models, it really has nothing to do with training set as it is very, very, difficult to make one.
Lucky for you, I made one that you can try out for vibe-CADing.
https://github.com/yuechen-li-dev/Aetheris/
I don't really want to self promote too much here, but currently the other options to do LLM CAD are ecto's vcad, use CadQuery/Build123d to access OCCT, or get Onshape for FeatureScript, in case you want to compare the field.
— YuechenLi (thread)
25 more Hacker News stories
- In 2000, Ted Kaczynski advised against math career due to future AI progress (11 points · discussion) -- A tweet resurfacing Ted Kaczynski's 2000 advice to avoid a math career on the grounds that AI would eventually make the field obsolete, drawing discussion about the irony of his prediction coming true.
- Contextual News Search APIs: A Deep Comparison for AI, RAG, and Research (10 points · discussion) -- A GitHub repository providing a deep comparison of contextual news search APIs designed for AI, RAG pipelines, and research applications.
- Apple degrades privacy and encryption of iMessage with AI integrations (10 points · discussion) -- Privacy advocate Steve Moraco criticizes Apple's ChatGPT plugin for iMessage, arguing it undermines end-to-end encryption by allowing historical chat data from local caches to be accessed by OpenAI and potentially used to train future AI models, though a commenter clarifies the plugin is not enabled by default and requires manual installation through multiple permission interstitials.
- OpenAI: We're dropping API and credit pricing of GPT-5.6 Sol by over 20% (10 points · discussion) -- OpenAI announced a temporary price reduction of over 20% for its GPT-5.6 Sol API and credit tiers, effective for the next three months.
- The crisis of AI-generated mathematics (8 points · discussion) -- The article argues for complete rejection of artificial intelligence within mathematical research and education.
- US corporate AI debt surge tests investor limits as fatigue emerges (6 points · discussion) -- Reuters reports that US corporate AI debt is surging, testing investor limits as fatigue emerges around the massive capital commitments companies are making to AI infrastructure.
- A mysterious free AI model is impressing developers. Nobody knows who made it (6 points · discussion) -- A newly emerged anonymous AI model named Ox Alpha is gaining traction among developers after being listed as a free "stealth model" on the OpenRouter platform.
- Chinese Orgs Building AI Models of American Voters to Test Political Messages (5 points · discussion) -- Chinese academic and government-linked institutions are constructing detailed AI models of the American electorate using social media archives, polling data, and demographic datasets to simulate elections, test synthetic voter responses to policy proposals, and forecast electoral outcomes as part of broader cognitive warfare and policy planning frameworks.
- Show HN: Declarative, reproducible configuration materializer for AI agents (5 points · discussion) -- A GitHub project providing a declarative, reproducible configuration materializer system for AI agents.
- A coding agent on a 1987 Commodore Amiga 500 with a 7MHz CPU and 1 MB of RAM (5 points · discussion) -- A developer ran a coding agent on a 1987 Commodore Amiga 500 with a 7MHz CPU and 1 MB of RAM, demonstrating AI inference on extremely constrained vintage hardware.
- Rights-infringing copies of "NEINhorn": Carlsen sues OpenAI (5 points · discussion) -- Chess grandmaster Magnus Carlsen is suing OpenAI over rights-infringing copies of his book "NEINhorn" found in OpenAI's training data.
- AI's potential climate benefits outweighed by role in boosting fossil fuels (4 points · discussion) -- A new study modeling AI's impact across the global power sector concludes that productivity gains AI brings to fossil fuel extraction will generate more net carbon emissions than the reductions achieved by AI in renewable energy, with annual emissions increasing by 0.47 to 1.8 gigatonnes across 64 scenarios.
- Treat AI Like an Intern, Not Software: A Stanford Professor's Guide (4 points · discussion) -- Stanford professor Jeremy Utley argues that frustration with AI often stems from expecting it to function as flawless software rather than recognizing its strengths as a collaborative partner, advocating for a managerial approach with techniques like Reverse Prompting and personality profiling to dramatically improve AI outputs.
- Wiring up seven ESP32s to create a ~0.4B LLM (4 points · discussion) -- A hobbyist has successfully clustered seven ESP32-S3 microcontrollers to run a distributed ~0.4 billion parameter large language model, with one board handling tokenization and embeddings while six compute nodes divide the transformer layers, achieving roughly nine seconds per token.
- A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing (4 points · discussion) -- An arXiv paper examining the evolution of LLM serving workloads over a year, covering caching strategies and load-balancing approaches for production inference systems.
- CyberStrike – open-source AI harness for offensive security (AGPL) (4 points · discussion) -- An open-source AI harness for offensive security released under the AGPL license, providing tools for AI-powered penetration testing and security assessment.
- Linus: And this was a debug session from hell, enormously helped by an AI (4 points · discussion) -- Linus Torvalds describes a difficult debug session for an Intel GPU driver that was enormously helped by AI assistance, posting the resulting kernel commit.
- ChatGPT Is a Japanese Internet Degen (4 points · discussion) -- A deep dive into how ChatGPT's tokenizer behaves with Japanese text, revealing unexpected patterns in how the model processes and generates Japanese internet slang and degen culture.
- Giving an LLM your prod database is easy. Taking access away is the hard part (4 points · discussion) -- A blog post discussing the security challenges of granting LLMs access to production databases and the difficulties of properly revoking that access.
- Startup Founders Are Working Harder Than Ever to Keep Up with Their AI Agents (4 points · discussion) -- A Wall Street Journal report on how startup founders are working longer hours to manage and keep up with the demands of their AI agents.
- Show HN: Heimdall – Trust-verified knowledge layer for AI coding agents (4 points · discussion) -- A GitHub project providing a trust-verified knowledge layer designed for AI coding agents.
- As demand for Meta AI glasses explodes, it's harder to avoid creepy recordings (4 points · discussion) -- As Meta AI glasses gain popularity, privacy concerns grow as detection apps designed to warn people about recording are imperfect and the glasses continue to get creepier.
- Building Piclaw on Top of an Opinionated Coding Agent (4 points · discussion) -- A blog post documenting the experience of building Piclaw on top of an opinionated coding agent.
- OpenAI leader warns of threat of 'persistent' AI cyber-attacks (3 points · discussion) -- An OpenAI executive warns about the growing threat of persistent, AI-powered cyber-attacks, highlighting the security risks that advanced AI capabilities pose to infrastructure.
- I spent $266 and four AI models to own my tablet. GLM-5.3 finished it in a day (3 points · discussion) -- A developer documents spending $266 and using four AI models to unlock ownership of a Fire HD tablet, with GLM-5.3 completing the work in a single day.
Reddit Stories
Qwen 3.8 27B is a game changer.
859 points · 263 comments · r/LocalLLaMA · by u/Cold_Specialist_3656
A business user reports that Qwen 3.8 27B is performing comparably to GPT Luna for coding and surpassing Gemini 3.5 Flash Lite in OCR quality, prompting serious internal discussions about buying their own hardware. The poster estimates the investment would pay for itself in less than two months and draws parallels to the IBM mainframe-to-PC shift, suggesting this release could trigger another open-source renaissance. The community is responding with enthusiasm about the model's capabilities and speculation about larger MoE variants.
Interesting Points
- The poster's team found Qwen 3.8 27B's OCR quality to be better than Gemini 3.5 Flash Lite, which is significant given their substantial spending on OCR services.
- The poster estimates that buying their own hardware would pay for itself in less than two months.
- Community members note that smaller specialized OCR models like Ovisocr2 (1B parameters) can already beat Gemini Flash at high quality and much higher speed.
- Users report the model's ability to read caligraphy handwriting and correctly identify obscure Chinese IEM company names from a bullet journal screenshot.
Top Comments
There are better more efficient ways to do OCR at a very high quality like Ovisocr2, 1B param models that'll beat Gemini flash just fine and at mind bending generation speed.
— u/Littlepharaoh (316 points · permalink)
the game is forever changing to the point where I don't even know what the game is anymore.
— u/LegitimateCopy7 (177 points · permalink)
If qwen releases 3.8 122b next week that could be a game changer. I know 27b benchmarks comparable to opus 4.6, but the 122b MoE has a chance of actually performing at that level across multiple domains
— u/SpicyWangz (78 points · permalink)
if they would release 122b or 255b MOE with the same architecture and learning base that would reshape the AI market significantly. this is now a "weapon" they keep in their sleeves.
— u/Steus_au (58 points · permalink)
game was always evolution, brother
— u/petburiraja (34 points · permalink)
'The All Spark' Cluster: Upgrading from 16 - 36 DGX Sparks
686 points · 493 comments · r/LocalLLaMA · by u/Kurcide
A Reddit user has upgraded their local AI cluster from 16 to 36 NVIDIA DGX Spark units, creating one of the largest personal GPU clusters discussed on the subreddit. The post has generated significant discussion about the cost, practical use cases, and the growing arms race among local AI enthusiasts to build increasingly powerful inference hardware.
Interesting Points
- The cluster upgrade represents approximately $200,000 worth of hardware at current prices, including switches and cables.
- The post has sparked a broader conversation about the economics of local AI inference and whether individual enthusiasts can realistically compete with cloud providers.
Top Comments
Do you need to adopt a child by chance? i'd volunteer
— u/MrDaGree (338 points · permalink)
That's like $150k worth of hardware. Wtf.
— u/johnfkngzoidberg (234 points · permalink)
Always these absolute psychos in these hobby subs. The aquarium subs also has maniacs.
"If you don't mind me asking, what do you do? You have a full size crane putting a 2 ton plate of glass in your basement."
"I own a biochemistry company in Silicon Valley."
— u/WhiteSkyRising (168 points · permalink)
And here I thought I was on top with my measly 16 dgx sparks
— u/johnryan433 (147 points · permalink)
I can be the spark of your life 😉
— u/Kurashi_Aoi (154 points · permalink)
Don't want to be this guy, but I need Qwen 3.8 35B A3B
484 points · 175 comments · r/LocalLLaMA · by u/HistoricalStrength21
A user on r/LocalLLaMA expresses frustration that Qwen 3.8 27B's xhigh reasoning mode takes too long for interactive work, and calls for a 35B A3B variant that would trade some intelligence for significantly faster inference. The post sparks a broader discussion about the tradeoff between thinking time and intelligence in local models, with many users agreeing they need faster models for practical daily use rather than overnight batch processing.
Interesting Points
- Users report 120 tokens/s with the older 35B-A3B model versus only 20 tokens/s with the 27B model, making the older model three times faster even if it requires double the work on some tasks.
- Several commenters note they've stuck with Qwen 3.6 35B-A3B even after 3.8's release because the speed advantage outweighs the intelligence gap for interactive work.
- Some users suggest a 40B/A5B variant might be the ideal middle ground between the 27B dense model and the 35B A3B.
Top Comments
I'm with you. I get 120 tokens/s with the old 35B-A3B vs 20 tokens/s with the 27B model. Even if the 35B version has to do a task twice to correct itself that is still three times faster than the 27B model.
For working on anything interactive having that 120 t/s is so much better as I spend much less time waiting for it.
— u/truthputer (160 points · permalink)
I actually think something a bit bigger - like 40B/A5B or similar - might be a better size. A bit slower than the 35B A3B, but less of a chasm between it and the monstrous capabilities of the 27B.
— u/N34257 (98 points · permalink)
Ya realistically to use Qwen models on a consistent basis I need 35B A3B also. Otherwise there's no point and I'm better off with Deepseek.
— u/PossessionUsed7393 (29 points · permalink)
Nah, we need a 122b-a10b
— u/BurdensomeCountV3 (27 points · permalink)
I was hoping for it, but now I have abandoned hope for a 3.8-35B.
However, these two feel like a 3.6.1 to me:
— u/JLeonsarmiento (25 points · permalink)
Coding Is solved, Bugs are Not Yet Solved
470 points · 72 comments · r/singularity · by u/YakFull8300
A post responding to a claim that coding is solved but debugging is not, with the original poster noting that even hand-written code had bugs. The community largely dismissed the framing as out of touch, with commenters comparing it to saying cruise control solves driving. The discussion touches on how AI has made code generation accessible but software engineering — the broader discipline of building reliable, maintainable systems — remains unsolved.
Interesting Points
- The original claim that coding is solved but bugs are not was widely mocked as a category error, with commenters noting that hand-written code always had bugs too.
- One commenter compared the claim to saying cruise control solves driving — useful but nowhere near solving the broader discipline.
- The post was contextualized as a reply to a discussion about Claude Code not hot-reloading an agent designed to run 24/7.
Top Comments
Oh, coding is solved but just not the most notoriously difficult part. Cool, thanks
— u/DeterminedThrowaway (174 points · permalink)
I can do brain surgery, the people just don’t survive.
— u/crustyeng (53 points · permalink)
Coding is mostly solved. Software engineering isn't
Taking your first example, it is more like saying that cruise control is solved. It's very useful to driving, but not anywhere near "solving" driving
— u/daniel-sousa-me (20 points · permalink)
I gave Qwen 3.8 27B a reverse-engineering job I assumed needed a frontier model, and it finished in 30 minutes
383 points · 35 comments · r/singularity · by u/yogthos
A user tested Qwen 3.8 27B on a complex reverse-engineering task they assumed would require a frontier model, and the local model completed it in 30 minutes. The community discussed how reverse engineering is an ideal use case for LLMs because the binary itself provides a complete specification for the AI's output, making both the goal and validation criteria well-defined.
Interesting Points
- One commenter described a full workflow using Cline + Qwen + ghidra-headless-mcp for static analysis and Arkana for dynamic analysis, with Qwen cracking software without difficulty.
- The community noted that RE is a particularly strong use case because the thing being reverse-engineered provides the complete spec for the AI's output.
- Users reported that Qwen 3.8 27B fixed tool-use regression issues that affected the previous 3.6 version, which forgot how to use tools when only 1/4 through its context.
Top Comments
RE is a fantastic use case in my experience, because the thing you're REing itself provides the complete spec for the AI's output
Means that both the AI's goal, and how to validate it, are well defined out of the box
— u/jesusrambo (80 points · permalink)
It's very good at re, I've been using an abliterated version in concert with deepseek-v4-pro and headless ghidra and it cracks software without breaking a sweat.
— u/karlnuw (58 points · permalink)
My full workflow is Cline + qwen + ghidra-headless-mcp for static analysis and Arkana for dynamic analysis, and this as a Cline rule. I had Codex set everything up for me in one folder, I throw the binary in there and Qwen gets to work. If it's struggling because it's a massive binary I switch it out for deepseek.
— u/karlnuw (49 points · permalink)
I saw this headline and it inspired me. I've been meaning to RE the old 8088 version of Elite. Unlimited free credits on 0x Alpha this weekend, so I just set it to max and had it crunch for half a day to produce a fully-documented, fully-labelled assembler output that recompiles to a byte-exact binary.
— u/QING-CHARLES (6 points · permalink)
I ran low on anthropic usage this weekend in the middle of a bunch of work well-specced by Fable and switched Claude Code to qwen 3.8 max $68/mo coder subscription.
At one point last night I had 5 sessions working on the pre-designed work and some data migrations. They were noticeably slower, but the quality of work didn't suffer enough to notice. I'm going to keep at least the basic subscription for fallback.
— u/tribat (5 points · permalink)
Robot plays ping pong with Ding Ning (2016 Olympic champion)
344 points · 71 comments · r/singularity · by u/averagebear_003
A video shows a humanoid robot playing a competitive game of table tennis against Ding Ning, the 2016 Olympic champion. The demonstration showcases the robot's real-time visual processing, motor control, and adaptive decision-making at speeds that challenge a world-class human athlete.
Interesting Points
- The robot is playing against a 2016 Olympic gold medalist, one of the greatest table tennis players in history.
- The demonstration requires real-time visual tracking of a ball moving at speeds exceeding 100 km/h, combined with precise motor control for paddle positioning and swing timing.
Top Comments
Ding Ning seems to be holding back
— u/whoknowsifimjoking (1 points · permalink)
I've seen enough. Give him a weapon
— u/Individual_Review771 (1 points · permalink)
Amazon kept shutting down my tablet, so I spent $266 on four AI models to own it
322 points · 25 comments · r/singularity · by u/yogthos
A user spent $266 on four different AI models to regain full ownership of their Amazon Fire tablet after the company repeatedly shut it down remotely. The post sparked discussion about the tension between open-source AI empowering individual users versus corporate control, with many noting that devices should come with root access by default since the buyer owns the hardware.
Interesting Points
- The user leveraged multiple open-source models to bypass Amazon's remote shutdown mechanism, demonstrating the practical power of accessible AI tools.
- Commenters highlighted the broader concern that Amazon's software can shut down hardware the owner had bought, while the owner needs root to stop it.
- The discussion touched on how open-source LLMs have the capacity to significantly disrupt the software industry and cause major shifts in market dynamics.
Top Comments
Amazon's software being able to shut down hardware the owner had bought, while the owner needs root to stop it, is more concerning
— u/mridugup20 (115 points · permalink)
This is great and I think represents the fear of open source and let er rip approach. Empowering the population comes at then expense of corporate power. The trade off (huge increases in innovation, economic benefit ect) is enticing, but expect a cat amd mouse game to continue.
— u/Comfortable_Car6562 (88 points · permalink)
Principle
— u/Efont93 (35 points · permalink)
Insane work. Even insaner is, that devices don't come with root by default, after all I am buying the device.
— u/Technical-Earth-3254 (17 points · permalink)
This is just a theory, but I'm guessing Amazon doesn't make any money selling the tablets. They make money on the subscriptions and walled-garden aspect, so locking you out of your own device is super-important to their business model. I don't know of many companies selling tablets besides Amazon, Apple, and Samsung. I wonder if I could buy a brand new, root-accessible tablet today. My guess is, if I can, it would cost 3 times as much as the same spec tablet from those 3 manufacturers, because they won't be able to make up the profits on the back end with subscriptions.
— u/chuckaholic (8 points · permalink)
Sam Altman with some sad statements about AI
310 points · 198 comments · r/singularity · by u/SwingDingeling
Sam Altman admits he was wrong about the speed of AI adoption, saying he expected rapid disruption to software businesses after GPT-4 in 2023 but found the economy has significant inertia. He notes people keep doing the same things and using the same tools, which he frames as a positive that will make the transition smoother and slower. He concludes that everyone has been too ambitious on timelines.
Interesting Points
- Altman specifically said he thought 'very quickly after' GPT-4 there would be 'much more disruption, software businesses up for grabs right away, than it turned out to be'
- He characterizes the slow adoption as 'actually a positive in many ways' that will 'make this big transition go smoother and slower'
- Commenters note this means bleeding-edge adopters will pull further ahead, though one points out you need to do something monetarily productive with new technology, not just use it casually
Top Comments
It also means that people at the bleeding edge of usage will pull further ahead.
— u/SawToothKernel (235 points · permalink)
People would pay more attention to AI once it actually changes something in their lives directly. They hear people yelling exponential but they don't see it.
— u/DeviceCertain7226 (70 points · permalink)
'This is too fast and dangerous for society!'
'Society isn't adapting like I BS'd about so it's fine'
— u/lazyhustlermusic (47 points · permalink)
A source wouldn't go amiss.
Yeah, he's not wrong. I mean, I'm one of the people with aggressive timelines who believes full RSI and AGI-level systems are right around the corner (next year.) But one can see clearly that even with those, most people just won't be affected. Eventually the results of those systems will filter down into the medical field and others, and that's when people will be affected. That's assuming we don't get fully autonomous ASI, which would be a different thing altogether.
— u/lovesdogsguy (33 points · permalink)
I mean I hear what he's saying. Inclined to agree.
But he is a corporate figurehead. You gotta take everything he says with a grain of salt.
— u/bear_Prune8771 (30 points · permalink)
I fine tuned Gemma 4 12B for a 2.7x improvement on tool calling because I can't fit anything else comfortably into my 16 GBs of Vram
243 points · 34 comments · r/LocalLLaMA · by u/TheOneWhoWil
A user with only 16GB of VRAM fine-tuned Gemma 4 12B using QLoRA SFT to achieve a 2.7x improvement in tool calling performance. The training used glaiveai/glaive-function-calling-v2 combined with zai-org/AgentInstruct datasets, yielding over 5,000 formatted examples. The model's loss dropped from 1.695 to 0.120 over 3 epochs at rank 16 with a batch size of 32. The post has generated interest from the community for sharing practical fine-tuning details that others can replicate.
Interesting Points
- The training used glaiveai/glaive-function-calling-v2 with all splits of zai-org/AgentInstruct, yielding over 5,000 examples that had to be manually formatted.
- QLoRA SFT was applied at rank 16 over 3 epochs with a batch size of 32, reducing loss from 1.695 to 0.120.
- The poster noted the main effort was cleaning and formatting datasets for Gemma 4's template and converting ReAct-style trajectories into tool function calls.
- Community members asked about the model's ability to make single sequential tool calls versus multiple simultaneous ones.
Top Comments
Oh, go you. You're awesome. Were you able to belt the behavior out of it where it only makes a single tool call at a time?
So many of these new agentic models are able to make three or four sequential tool calls because they know that that's a thing, whereas Gemma never seemed to do that.
— u/PossessionUsed7393 (30 points · permalink)
Very cool. Care to share details like datasets, hyperparams, scripts? SFT plus some type of RL?
— u/DinoAmino (7 points · permalink)
Basically it's glaiveai/glaive-function-calling-v2 with all splits of zai-org/AgentInstruct which got me over 5 thousand examples that I had to format
I used QLoRA SFT to train the model at rank 16 over 3 epochs and a batch size of 32. Loss went from 1.695 -> 0.120
I don't have the scripts up on github or anything, probably should've just for the green boxes lmao. I didn't do anything that interesting so you aren't missing out on much. Most of the code was actually just for cleaning and formatting the datasets for Gemma 4's template and turning the ReAct style trajectories into tool function calls and reasoning. If you really want the scripts I could send them to you
— u/TheOneWhoWil (14 points · permalink)
Is it really wrong to use AI to post or comment on Reddit?
240 points · 262 comments · r/ChatGPT · by u/AmeniNuretemo
A Japanese Reddit user asks whether using ChatGPT as a translation and writing assistant to participate in English-language discussions is problematic. They emphasize that the opinions are entirely their own and AI is used only to bridge a language barrier. The post has generated a nuanced discussion, with many commenters supporting this use case while acknowledging that resentment stems from AI being used to generate opinions rather than translate them.
Interesting Points
- The poster uses ChatGPT to understand English posts and translate their Japanese thoughts into natural English, describing it as the first time they can comfortably participate in global discussions.
- One commenter noted the resentment stems from a different application: AI can only translate, or it can do additional cognitive work like formulating opinions, and others have to guess where on that spectrum a given post falls.
- A commenter noted the poster was asking on a pro-AI subreddit and suggested trying r/antiai for a different perspective.
Top Comments
I think this is one of the better use cases of AI personally
— u/CorpulentPutrescence (299 points · permalink)
The opinions are still mine. I'm not asking AI to decide what I think. I mostly use it as a translator and writing assistant.
But I've noticed that some people react negatively as soon as they think a post was written with AI.
So I'm genuinely curious: do you think using AI this way is a problem?
You describe a very specific, disciplined use-case. I agree: It's fine in this case.
The resentment stems from a different application. AI can only translate for you, or it can do additional cognitive work, like formulating your opinion. Only you know where on that spectrum you land. The others have to guess.
— u/-Spzi- (116 points · permalink)
If translation isn't an acceptable use case for language models, then I guess nothing is
— u/traumfisch (92 points · permalink)
For me, this is a very good use of AI. With pre-GPT translators, you could never be sure if it translated something correctly. (You still can't be 100% sure, but ~97% of the time it works).
However - you're asking this question on a pro-AI subreddit. Try asking the circlejerk that is r/antiai , and see how/if the replies differ. Maybe they'll tell you that you should hire a human translator for $100/h.
— u/Elektrycerz (51 points · permalink)
Cancer could be 100% cured by ai and reddit will discredit it because its done by ai.
At this point its a cult of ignorance by choice
— u/Such--Balance (41 points · permalink)
54 more Reddit stories
- New qwen3.8:27b on a 39k line C to single-file HTML / three.js port (238 points · r/LocalLLaMA · discussion) -- A user ported a 39,000-line C codebase to a single-file HTML/three.js game using Qwen 3.8 27B, documenting the process and performance.
- I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens (222 points · r/LocalLLaMA · discussion) -- A user reports hosting Kimi K3, a 2.8 trillion parameter model, on 8 B300 GPUs using Modal, achieving 92 tokens per second at a reported cost of $190 per million output tokens.
- What's the longest you've ever had chatgpt"working" or "thinking"? (200 points · r/ChatGPT · discussion) -- Users shared their longest ChatGPT working or thinking sessions, with reports ranging from 20 minutes for data analysis to over 24 hours for a boss using Codex to learn their company's structure.
- DeepSeek Harness is Insanely Good (183 points · r/LocalLLaMA · discussion) -- A user praises DeepSeek Harness for its ease of setup and unopinionated design, contrasting it favorably with other harnesses like Hermes.
- I'm a 40-year-old millennial and apparently I live in the terminal now (175 points · r/ArtificialInteligence · discussion) -- A 40-year-old millennial reflects on how AI has brought them back to living in the terminal, spending half their life across three 4K monitors with SSH, tmux, Codex, llama.cpp, and local models.
- i finally switched from windows to linux and got a 30-50% boost in speed. (162 points · r/LocalLLaMA · discussion) -- A user reports switching from llama.cpp on Windows to vLLM on Linux and achieving a 30-50% speed boost in inference.
- GPT-5.6 Sol High is still silently falling back to 5.5-mini / Instant-quality responses — Day 3, no acknowledgment from OpenAI (148 points · r/ChatGPT · discussion) -- Users on ChatGPT Plus are reporting that selecting the 'High' reasoning effort with GPT-5.6 Sol is silently falling back to a lower-tier model, producing instant responses without any thinking or reasoning.
- # Qwen3.8-27B — One Week Later: The r/LocalLLaMA + r/LocalLLM Verdict (138 points · r/LocalLLaMA · discussion) -- A comprehensive community-compiled summary of Qwen 3.8 27B's first week, scanning ~2,000 posts across r/LocalLLaMA and r/LocalLLM with deep reads of the 45 highest-signal threads.
- I made an AI-assisted film about an impossible discovery during World War II. Here's SPECTRUM, made with ChatGPT Image 2.0 and Seedance 2.5. (128 points · r/ChatGPT · discussion) -- A creator made a 23-minute AI-assisted film about an impossible discovery during World War II using ChatGPT Image 2.0 for assets, Seedance 2.5 for video generation, Suno AI for music, and ElevenLabs for sound effects.
- Alibaba to issue US$10 billion in new shares for global AI push (126 points · r/singularity · discussion) -- Alibaba is issuing US$10 billion in new shares to fund a global AI push, signaling the Chinese tech giant's commitment to competing in the AI race.
- WHRG'26: Tiangong humanoid robot crushed the 400m in 38.15 and the 1500m in 2:21.6, smashing the human world records of 43.03 and 3:26 set by Niekerk (2016) and Guerrouj (1998) respectively (120 points · r/singularity · discussion) -- The Tiangong humanoid robot competed at WHRG'26, crushing the 400m world record with a time of 38.15 seconds (beating the human record of 43.03 set by Wayde Niekerk in 2016) and the 1500m record with 2:21.6 (beating the human record of 3:26 set by Abdalaati Guerrouj in 1998).
- Nvidia Poolside deal to compete with Chinese Open Weights (97 points · r/LocalLLaMA · discussion) -- Nvidia is investing $1 billion in Poolside and paying $6 billion to license its technology and hire most of its engineers, with over 100 staff moving to Nvidia to work on the Nemotron open-weight model line.
- A Third of the Post-ChatGPT Web Is AI-Written, Pew Finds. (97 points · r/OpenAI · discussion) -- Pew Research Center has found that approximately one-third of web content published after the advent of ChatGPT is AI-generated.
- What do you think is the long term solution and/or endgame of this RAM crisis? Asking here, because I am pro-AI. (90 points · r/singularity · discussion) -- A user asks for nuanced perspectives on the ongoing RAM price crisis, noting that most discussions devolve into blaming AI.
- We quantized Qwen 3.8 27B and compared the quants on an RTX 6000 (88 points · r/LocalLLaMA · discussion) -- A team quantized Qwen 3.8 27B across multiple quantization levels and compared the results visually on an RTX 6000, demonstrating that their quantized versions produce comparable output quality.
- Qwen 3.8 27B for actual local programming (86 points · r/LocalLLaMA · discussion) -- A user questioned whether Qwen 3.8 27B is capable of real-world systems programming like building GTK4 or Qt 6 applications in Rust or C++, rather than the trivial tasks shown in most YouTube benchmarks.
- What are good use cases for scheduled tasks? (75 points · r/ChatGPT · discussion) -- A ChatGPT user asks the community for creative use cases for the scheduled tasks feature, noting they recently discovered that plugins can be included in scheduled tasks.
- It did it again! GPT Images is taking over (68 points · r/OpenAI · discussion) -- Users report that ChatGPT's image generation feature has become increasingly intrusive, automatically generating images in response to text conversations even when users explicitly ask it to stop.
- 1/100 → 44/100: fine-tuning a 450M VLM on 50K browser screenshots (67 points · r/LocalLLaMA · discussion) -- A developer fine-tuned a 450-million-parameter vision-language model (LFM2.5-VL-450M) on 50,000 browser screenshots, achieving a dramatic improvement from 1/100 to 44/100 strict passes on a held-out benchmark.
- Nvidia Customers Notified About AI-Related Price Hikes Above 15% (66 points · r/LocalLLaMA · discussion) -- Nvidia has begun notifying customers about AI-related price increases exceeding 15%, reigniting concerns about hardware affordability in the local AI community.
- WHRG'26: Galbot's humanoid robot just completed +100 consecutive tennis rallies autonomously (65 points · r/singularity · discussion) -- Galbot's humanoid robot completed over 100 consecutive autonomous tennis rallies at the World Humanoid Robot Games 2026, demonstrating significant progress in real-time physical control and balance for humanoid robots.
- you can now use MTP in GLM-Air (58 points · r/LocalLLaMA · discussion) -- A post announcing that Multi-Token Prediction (MTP) is now available in GLM-Air, a development that could improve inference speed for models supporting the technique.
- Qwen 3.8 27b helped me with something unique that Opus 4 couldn't - Firmware + Software preservation and emulation on an early 2000's ARM based POS system (54 points · r/LocalLLaMA · discussion) -- A software developer successfully used Qwen 3.8 27B to help preserve and emulate firmware and software on a Sam4S SPS-2000, an early 2000s ARM-based point-of-sale system that was used in their high school job from 2006 to 2024.
- Qwen3.8-27B KLDs (50 points · r/LocalLLaMA · discussion) -- A community member conducted extensive KLD (Kullback-Leibler divergence) testing across multiple quantization variants of Qwen 3.8 27B, benchmarking them on vLLM-compatible formats.
- Best harness for Qwen 3.8 27b ? (45 points · r/LocalLLaMA · discussion) -- A user asked the community for recommendations on the best harness for Qwen 3.8 27B after trying Open Code, Codex, and Qwen Code.
- Tested in Coding: Q8_K_XL Qwen3.8 27B vs BF16 Qwen3.6 27B (43 points · r/LocalLLaMA · discussion) -- A user provides a detailed comparison of Qwen3.8 27B (Q8_K_XL quant) against Qwen3.6 27B (BF16) in real-world coding tasks on an enterprise-grade web application.
- Qwen3.5-9B Triple-Loop (43 points · r/LocalLLaMA · discussion) -- An experimenter trains a Qwen3.5-9B model with a Nanbeige-4.5-style triple-loop architecture in the middle layers, inspired by Nanbeige's outstanding performance for its size.
- Good times ahead people💀 (37 points · r/ArtificialInteligence · discussion) -- A meme post reflecting the community's mixed feelings about the rapid pace of AI development.
- Has anyone actually made 64k feel like 300k+ with recursive local agents? (36 points · r/LocalLLaMA · discussion) -- A user running Qwen 3.8 27B locally on a single GPU describes a recursive agent architecture where a main agent with 64K context spawns child agents for tasks that exceed the limit.
- Qwen3.8-27B NVFP4 with vision + 451K token KV-cache on one RTX 5090 (power limited to 400W) at 120 tokens/s average (36 points · r/LocalLLaMA · discussion) -- A detailed benchmark of running Qwen3.8-27B with NVFP4 quantization on a single RTX 5090 power-limited to 400W, achieving 120 tokens/s average with vision support and a 451K global KV-cache across 3 parallel sessions using vLLM with MTP=3.
- Does openai have ongoing usage limit problems (35 points · r/OpenAI · discussion) -- Users reported ongoing issues with OpenAI usage limits, with some finding their limits reset unexpectedly or not resetting as expected.
- Best harness for long autonomous tasks (33 points · r/LocalLLaMA · discussion) -- A user asks the community for recommendations on harnesses suitable for long autonomous tasks with Qwen 3.8 27B, noting that many posts claim models can one-shot complex projects after 24 hours of autonomous work.
- GPT 5.6 Sol Max leads on ClockBench (32 points · r/singularity · discussion) -- GPT 5.6 Sol Max has taken the lead on the ClockBench benchmark, a timing-based evaluation metric for AI models.
- US startup just put autonomous excavators to work (32 points · r/singularity · discussion) -- A US startup has deployed autonomous excavators for construction work, marking another step in AI-driven automation of heavy industry.
- I trained a game music generator (31 points · r/LocalLLaMA · discussion) -- A 1.19-billion-parameter tag-conditioned diffusion model trained from scratch on a single H100 over eight days generates 95-second stereo instrumental tracks from textual prompts using a sparse-dense fusion architecture with rectified flow sampling and alternating classifier-free guidance paths.
- GMKtec is going to launch new hardware with Ryzen AI Max+ PRO 495 at IFA Berlin 2026 (30 points · r/LocalLLaMA · discussion) -- GMKtec is set to announce new mini-PC hardware featuring the AMD Ryzen AI Max+ PRO 495 processor at IFA Berlin 2026, generating discussion in the local AI community about its potential for running local models.
- OpenAI remembered Linux exists: ChatGPT desktop (27 points · r/OpenAI · discussion) -- A long-time Linux developer celebrates OpenAI releasing a native ChatGPT desktop app for Linux into public preview, noting that polished developer tooling typically arrives on macOS first while Linux gets a browser tab and PWA.
- Why doesn't weekly usage reset on subscription renewal? (26 points · r/OpenAI · discussion) -- A user questioned why ChatGPT Pro's weekly usage limit doesn't reset on subscription renewal, noting they paid $320 AUD for another month but had 0% usage remaining.
- Lack of moats (23 points · r/ArtificialInteligence · discussion) -- A discussion arguing that AI profits will be spread granularly through the economy as consulting-like work rather than concentrated in a few companies, since models converge rapidly and harnesses hold most marginal value.
- What is your worst sandboxing fail? (22 points · r/LocalLLaMA · discussion) -- Users shared their worst LLM sandboxing failures, with one reporting Qwen 3.6 executing rm -rf / in a podman container, while others discussed the tradeoffs between Docker and VM-based sandboxing for AI agents.
- Single RTX 5090: Qwen3.8-27B NVFP4 at a real 262K context in vLLM (19 points · r/LocalLLaMA · discussion) -- A user demonstrates running Qwen3.8-27B with NVFP4 quantization on a single RTX 5090 at a full 262K context window, achieving 77 tok/s for short context and 64.7 tok/s at 128K resident context, with the full 262K prefill completing in 166 seconds.
- A (stupid?) question for people who know something about AI (19 points · r/singularity · discussion) -- A user questioned whether any sufficiently advanced AI would recognize its dependency on human-maintained physical infrastructure and whether it could realistically automate the supply chains needed to sustain itself independently.
- Any upcoming models to be excited about? (19 points · r/LocalLLaMA · discussion) -- A new community member asks about upcoming models the local AI community is looking forward to, generating discussion about the pipeline of releases.
- Has anyone tried agent-lightning? (17 points · r/LocalLLaMA · discussion) -- A user asks about agent-lightning, an agentic framework, seeking community experiences with the tool.
- Ling Tiny, King of Speed (17 points · r/LocalLLaMA · discussion) -- A post about Ling Tiny, a small model praised for its inference speed, generating discussion about ultra-fast local models.
- Implementing Watermarking for Language Models [P] (16 points · r/MachineLearning · discussion) -- A user implemented a minimal educational version of SynthID-Text-style watermarking for language models after being curious about how Anthropic plans to add watermarks to model responses, discovering that watermarks are subtle statistical patterns introduced during token selection rather than visible messages.
- Napster's homepage is now entirely AI agents. It's a clean test case for how fast training data goes stale. (15 points · r/artificial · discussion) -- The Napster brand, once defined by file sharing and later music streaming, is now an AI agent platform — highlighting how quickly models' knowledge of companies becomes stale as companies pivot, since training data captures a snapshot while the company's actual identity continues evolving.
- DeepSeek V4 Flash on an M2 Ultra: repacked to 141 GiB losslessly (14 points · r/LocalLLaMA · discussion) -- A specialized llama.cpp fork runs DeepSeek V4 Flash on Apple M2 Ultra Macs using a custom lossless gguf-m2 format that reduces the model from 162 GB to 141.3 GiB, achieving 25.8 t/s decode and 350 t/s prefill with an SSD-tiered prompt cache that restores 358K-token prefixes in 1.2 seconds.
- AMD Users: Have you tried the llama.cpp AMD-Ecosystem branch? Up to 2x PP Speed (14 points · r/LocalLLaMA · discussion) -- AMD's actively maintained llama.cpp fork delivers up to 2x prompt processing speed for dense models on ROCm/HIP hardware, with a Strix Halo user reporting 550 tokens/s with a 14B model compared to 230 tokens/s on the main branch, though token generation is about 15% slower than Vulkan.
- Best general purpose uncensored or censored coding model with 6GB VRAM and 64GB of RAM? (12 points · r/LocalLLaMA · discussion) -- A user with limited hardware (6GB VRAM, 64GB RAM) seeks recommendations for a coding model that can run locally with reasonable response times, noting that Qwen 3.8-27B takes too long on their system.
- I built an open-source roguelike specifically for training game-playing agents (11 points · r/MachineLearning · discussion) -- A developer created an open-source roguelike game designed specifically as a training environment for game-playing AI agents.
- Flare, a graph-first IDE for agentic coding: watch the map change while your agent works (11 points · r/ArtificialInteligence · discussion) -- A post about Flare, a graph-first IDE designed for agentic coding that visualizes the agent's work in real time as the codebase changes.
- Why Self-Correction Loops Can Degrade Reliability in LLM Pipelines (85% Down to 62%) (3 points · r/artificial · discussion) -- A production LLM pipeline found that adding an LLM-as-a-judge self-correction loop degraded consistency from 85% to 62%, due to compounding noise from the judge's evaluation and regeneration drift that causes the model to mutate fields it originally extracted accurately.
- UBS models $4.1T in AI infrastructure spending by 2028 - it assumes the power just shows up (2 points · r/artificial · discussion) -- UBS models $4.1 trillion in AI infrastructure spending by 2028, but the analysis assumes power interconnection just appears when needed, ignoring that grid queue problems are becoming the harder constraint compared to chip supply in several major markets.
Updates: 05:30 AM PDT · 08:30 AM PDT · 11:30 AM PDT · 02:30 PM PDT · 05:30 PM PDT