· 05:30 PM PDT

Qwen 3.8 Dominates Local AI as Humanoid Robots Shatter Records

Overview

Qwen 3.8 27B has taken over developer forums, delivering frontier-level coding and reasoning capabilities on modest local hardware while sparking intense debate over model quantization and VRAM optimization. This surge in accessible AI is matched by rapid progress in robotics, with humanoid machines recently shattering human world records in sprinting and autonomous tennis. Meanwhile, industry executives are recalibrating expectations around AI adoption speed and enterprise data dependency, even as major labs adjust pricing and release new desktop tools. The day’s conversation underscores a clear pivot: highly capable AI is no longer confined to massive cloud clusters, and its physical applications are accelerating faster than anticipated.


Hacker News Stories

'AI refuser' quit her dream job, and hopes others follow

34 points · 39 comments · by mcapodici

The AFL logo

Gabrielle Boyle, a former AFL participation manager, resigned after the league refused to exempt her from using Microsoft Copilot AI, highlighting a significant gap in Australian workplace law. Current employment and privacy regulations generally allow employers to mandate AI tools on company equipment without offering workers an opt-out. While unions argue the Fair Work Act requires meaningful consultation when AI changes job functions, legal experts stress that individual exemptions remain unenforceable and ineffective at preventing broader data collection. The case illustrates how poorly managed AI implementations are driving away skeptical staff and exposing urgent needs for updated governance frameworks.

Interesting Points
  • An internal survey revealed 35% of AFL staff were uncomfortable with AI adoption, yet more than half of Australian Services Union clerical workers were unaware of their employer's AI policy entirely.
  • The AFL's official AI policy explicitly endorses ChatGPT, Copilot, and Claude Code, while a footnote admits the policy document itself was drafted with generative AI assistance.
  • New South Wales recently passed the Work Health and Safety Amendment (Digital Work Systems) Act, though its application to general-purpose assistants like Copilot has not been legally tested.
  • Research by KPMG and the University of Melbourne shows Australians rank among the least trusting of AI globally, with skepticism driven primarily by employer conduct and change management.
Top Comments

I really hate how the excruciatingly overdue reckoning for tech elites and income inequality is turning into this religious schism over AI and datacenters.

We're going to smash the looms rather than figure out how to share them, again.

cagenut (3 replies)

The underlying issue is really capital ownership, what classifies as capital, protections around it, and how much leverage capital provides on a society and democracy (that part being the most important, IMHO).

Who owns the looms and what they do with them isn’t inherently an issue if all loom owners can do is buy an extra Yacht. Instead they can enforce ungoverned law on society through a combination of disproportionate influence in government and through market forces where private policy (especially at large) become nearly undifferentiated from law (the policy that benefits them becomes so widespread and normalized that alternatives are for all intent and purposes impractical or unreasonable, therefor private policy within a capital ownership domain is law or they’re a monopoly so their policy is the policy).

But I think you’re right, as always we’re going to focus on the adjacent issues vs addressing the root of the problem. The issue is what wealth inequality affords one, when it is capable of infringing on rights and livelihoods of others, not that they necessarily have to share the other luxuries and rewards of their attained wealth. I care not how many luxuries in life Musk has, I may be envious from time to time but whatever. I care a lot more when things he does or says has unrealistic influence and affects me directly, just because he sits atop a mountain of capital and we pretend that mountain of capital somehow was bestowed upon him from divinity that he should have such influence. I’m picking on Musk because he’s the richest and has clear examples of this, he’s by no means alone… it’s that class of wealth at large.

Frost1x (0 replies)

I don't appreciate the celebration around acts of individualism rather than collectivizing to improve our workplace for all.

wildrhythms (3 replies)

AI boosters on this site will be very mad that this person exists, but I'm glad there are people that stick to their principles. Principles that aren't just "productivity" or "efficiency". People can and should have other motives that are just as if not more important.

tkel (2 replies)

Does Australia have privacy provisions similar to the gdpr?

CalRobert (2 replies)


Software Engineering in the Agentic Era

26 points · 9 comments · by silverpiranha

Software Engineering in the Agentic Era

Simon Willison is launching "Agentic Engineering Patterns," a continuously updated guide documenting best practices for professional software engineers using autonomous coding agents. He defines "agentic engineering" as a discipline where engineers leverage tools that can both generate and execute code to test and iterate independently, contrasting it with "vibe coding." The project is structured as a series of evergreen chapters inspired by classic software design patterns, with the first two already published. Willison plans to release one to two new chapters weekly while maintaining a strict policy that all written content remains human-authored.

Interesting Points
  • The first two published chapters examine how near-zero initial code generation costs disrupt traditional engineering intuitions and how red/green test-driven development helps agents write more succinct, reliable code with minimal prompting.
  • Willison's existing archive of AI-assisted programming posts has already surpassed 345 entries, highlighting the need for a centralized, structured resource.
  • The Django backend and views powering the new guide format were almost entirely written by Claude Opus 4.6 running in Claude Code via an iPhone.
  • Unlike standard blog posts, the guide's chapters are designed as evergreen content that is updated over time rather than being frozen at first publication.
Top Comments

A few nits:

  1. Writing code is cheap now

Change to "generating code is cheap", reserve "writing" for the manually written code for the pre-AI era. I think this is a good wording separation.

  1. Agentic Engineering Patterns

I must ask to add at least a chapter to be read by agent. I.e., patterns just to tell agents how human might be working when working with them. Without this, I believe the book's content will be less relevant in 3 months, but with that, it feels a agentic-native book to me. (this is not try to be cute, we have to write for agents now)

bigcat12345678 (0 replies)

one theory I am thinking through is that people say AI is good at greenfield and has a harder time editing code in a legacy code base.

So maybe you just treat projects as greenfield instead of editing.

Like if you want to make changes to a page you generate a new version of the page instead of editing the old one in place.

Then you keep the old version as fallback if problems come up with the new version.

pianopatrick (1 reply)


Palantir's Karp – frontier AI labs that are 'trying to drug addict us'

19 points · 8 comments · by rishabhd

Palantir's Karp – frontier AI labs that are 'trying to drug addict us'

Palantir CEO Alex Karp warned that frontier AI labs are creating dependency among enterprises, arguing that companies need to protect their data or risk losing their business to model makers. He drew a parallel to Chinese models distilling U.S. models, noting that frontier labs have already distilled all the value of intellectual property everywhere. The comments suggest the argument, while potentially valid, is self-serving since Palantir's own forward deployed engineering and consultancy model faces the same threat from frontier models.

Interesting Points
  • Karp argued that enterprises need to protect their data or risk losing their business to model makers.
  • He said Chinese models can't be blamed for distilling U.S. models when the frontier labs "distilled all the value of IP, everywhere."
  • The comments suggest the argument, while potentially valid, is self-serving since Palantir's own forward deployed engineering and consultancy model faces the same threat from frontier models.
Top Comments

*"Karp said enterprises need to protect their data or risk losing their business to model makers.

Karp said Chinese models can’t be blamed for distilling U.S. models when the frontier labs “distilled all the value of IP, everywhere.”*

I think all of us would echo that.

andsoitis (2 replies)

While the argument makes sense, the motivation is rather self-serving: he is afraid frontier models will steal his forward deployed engineering/consultancy model.

Palantir is stealing business from consultancy firms, now frontier models are going to do the same to them. He has to convince executives (main purchase decision makers) by what he does best: politics.

mgh2 (0 replies)

Sauron warns against Saruman.

andrewstuart (0 replies)


Andrew Ng: AI Engineering Skills Map: Building and Deploying AI Applications

15 points · 0 comments · by Anon84

Andrew Ng AI Engineering Skills Map infographic

Andrew Ng expands on his AI Engineering Skills Map by detailing the core competencies required for building and deploying AI applications. He argues that the fundamental challenge of AI engineering lies in the inherent unpredictability of model outputs, which transforms development into a highly iterative, experimental process rather than a linear one. To navigate this uncertainty effectively, Ng outlines six essential technical domains: LLM foundations, data grounding, agentic system design, evaluation-driven development, production operations, and machine learning fundamentals.

Interesting Points
  • The skills map was constructed by analyzing a large volume of job postings, conducting structured expert interviews, and reviewing survey responses.
  • Understanding LLM internals like tokenization and context windows helps engineers decide when to rely on models, manage tradeoffs in knowledge cutoffs, and optimize for cache hits and sampling parameters.
  • Beyond early vector search RAG, grounding techniques now include knowledge graphs and semantic layers over structured data, requiring engineers to decide what to hardcode in prompts versus what to retrieve on-demand via tools.
  • Ng identifies driving a disciplined evals and error analysis loop as the single most distinguishing trait of top AI engineers, emphasizing the need to blend deterministic code-based checks, LLM-as-a-judge methods, and human-in-the-loop feedback.
  • Operating AI in production demands statistical regression testing calibrated to risk, alongside continuous monitoring for model drift, adversarial prompt injections, and cost/latency optimization through techniques like model distillation.

Why can AI generate Super Mario but not a wedge ramp for my robot vacuum?

11 points · 5 comments · by zhuchaokn

A discussion on Hacker News about the surprising gap between AI's ability to generate recognizable cultural artifacts like Super Mario and its difficulty with practical engineering tasks like designing a wedge ramp for a robot vacuum. Commenters suggest that parametric design tools like OpenSCAD, combined with LLMs, can produce surprisingly good results for part design when given metric units and sketches as context.

Interesting Points
  • One commenter reports getting Claude to generate a complex mating part from a 3D scan using OpenSCAD, with back-and-forth refinement
  • Another commenter points out that the real challenge isn't modeling but decomposing problems into usable chunks and anticipating what will be a problem
  • A developer shares Aetheris, an open-source geometric modeling kernel for LLM-based CAD, noting it competes with ecto's vcad and CadQuery/Build123d approaches
Top Comments

Ive had great success with asking claude to use openscad for part design.

Because it's parametric design as opposed to modeling, it inherently lends itself to making it easier to build.

Use metric units when prompting, and Ive even drawn sketches to add to context, and be surprised at how good the outcome can be!

chews (thread)

modeling isn't the hard part. It's decomposing it into usable chunks, the knowledge of what's going to be a problem, etc.

either way, give an actual CAD program like fusion a go. No, not the FOSS stuff- the real stuff. Although, some of them are "close, but no cigar" these days. Not like gimp which is an exercise in self-flagellation.

butvacuum (thread)

Well, because you need a geometric modeling kernel to generate BRep CAD models, it really has nothing to do with training set as it is very, very, difficult to make one.

Lucky for you, I made one that you can try out for vibe-CADing.

https://github.com/yuechen-li-dev/Aetheris/

I don't really want to self promote too much here, but currently the other options to do LLM CAD are ecto's vcad, use CadQuery/Build123d to access OCCT, or get Onshape for FeatureScript, in case you want to compare the field.

YuechenLi (thread)


25 more Hacker News stories

Reddit Stories

Qwen 3.8 27B is a game changer.

859 points · 263 comments · r/LocalLLaMA · by u/Cold_Specialist_3656

A business user reports that Qwen 3.8 27B is performing comparably to GPT Luna for coding and surpassing Gemini 3.5 Flash Lite in OCR quality, prompting serious internal discussions about buying their own hardware. The poster estimates the investment would pay for itself in less than two months and draws parallels to the IBM mainframe-to-PC shift, suggesting this release could trigger another open-source renaissance. The community is responding with enthusiasm about the model's capabilities and speculation about larger MoE variants.

Interesting Points
  • The poster's team found Qwen 3.8 27B's OCR quality to be better than Gemini 3.5 Flash Lite, which is significant given their substantial spending on OCR services.
  • The poster estimates that buying their own hardware would pay for itself in less than two months.
  • Community members note that smaller specialized OCR models like Ovisocr2 (1B parameters) can already beat Gemini Flash at high quality and much higher speed.
  • Users report the model's ability to read caligraphy handwriting and correctly identify obscure Chinese IEM company names from a bullet journal screenshot.
Top Comments

There are better more efficient ways to do OCR at a very high quality like Ovisocr2, 1B param models that'll beat Gemini flash just fine and at mind bending generation speed.

u/Littlepharaoh (316 points · permalink)

the game is forever changing to the point where I don't even know what the game is anymore.

u/LegitimateCopy7 (177 points · permalink)

If qwen releases 3.8 122b next week that could be a game changer. I know 27b benchmarks comparable to opus 4.6, but the 122b MoE has a chance of actually performing at that level across multiple domains

u/SpicyWangz (78 points · permalink)

if they would release 122b or 255b MOE with the same architecture and learning base that would reshape the AI market significantly. this is now a "weapon" they keep in their sleeves.

u/Steus_au (58 points · permalink)

game was always evolution, brother

u/petburiraja (34 points · permalink)


'The All Spark' Cluster: Upgrading from 16 - 36 DGX Sparks

686 points · 493 comments · r/LocalLLaMA · by u/Kurcide

'The All Spark' Cluster: Upgrading from 16 - 36 DGX Sparks

A Reddit user has upgraded their local AI cluster from 16 to 36 NVIDIA DGX Spark units, creating one of the largest personal GPU clusters discussed on the subreddit. The post has generated significant discussion about the cost, practical use cases, and the growing arms race among local AI enthusiasts to build increasingly powerful inference hardware.

Interesting Points
  • The cluster upgrade represents approximately $200,000 worth of hardware at current prices, including switches and cables.
  • The post has sparked a broader conversation about the economics of local AI inference and whether individual enthusiasts can realistically compete with cloud providers.
Top Comments

Do you need to adopt a child by chance? i'd volunteer

u/MrDaGree (338 points · permalink)

That's like $150k worth of hardware. Wtf.

u/johnfkngzoidberg (234 points · permalink)

Always these absolute psychos in these hobby subs. The aquarium subs also has maniacs.

"If you don't mind me asking, what do you do? You have a full size crane putting a 2 ton plate of glass in your basement."

"I own a biochemistry company in Silicon Valley."

u/WhiteSkyRising (168 points · permalink)

And here I thought I was on top with my measly 16 dgx sparks

u/johnryan433 (147 points · permalink)

I can be the spark of your life 😉

u/Kurashi_Aoi (154 points · permalink)


Don't want to be this guy, but I need Qwen 3.8 35B A3B

484 points · 175 comments · r/LocalLLaMA · by u/HistoricalStrength21

A user on r/LocalLLaMA expresses frustration that Qwen 3.8 27B's xhigh reasoning mode takes too long for interactive work, and calls for a 35B A3B variant that would trade some intelligence for significantly faster inference. The post sparks a broader discussion about the tradeoff between thinking time and intelligence in local models, with many users agreeing they need faster models for practical daily use rather than overnight batch processing.

Interesting Points
  • Users report 120 tokens/s with the older 35B-A3B model versus only 20 tokens/s with the 27B model, making the older model three times faster even if it requires double the work on some tasks.
  • Several commenters note they've stuck with Qwen 3.6 35B-A3B even after 3.8's release because the speed advantage outweighs the intelligence gap for interactive work.
  • Some users suggest a 40B/A5B variant might be the ideal middle ground between the 27B dense model and the 35B A3B.
Top Comments

I'm with you. I get 120 tokens/s with the old 35B-A3B vs 20 tokens/s with the 27B model. Even if the 35B version has to do a task twice to correct itself that is still three times faster than the 27B model.

For working on anything interactive having that 120 t/s is so much better as I spend much less time waiting for it.

u/truthputer (160 points · permalink)

I actually think something a bit bigger - like 40B/A5B or similar - might be a better size. A bit slower than the 35B A3B, but less of a chasm between it and the monstrous capabilities of the 27B.

u/N34257 (98 points · permalink)

Ya realistically to use Qwen models on a consistent basis I need 35B A3B also. Otherwise there's no point and I'm better off with Deepseek.

u/PossessionUsed7393 (29 points · permalink)

Nah, we need a 122b-a10b

u/BurdensomeCountV3 (27 points · permalink)

I was hoping for it, but now I have abandoned hope for a 3.8-35B.

However, these two feel like a 3.6.1 to me:

https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B

u/JLeonsarmiento (25 points · permalink)


Coding Is solved, Bugs are Not Yet Solved

470 points · 72 comments · r/singularity · by u/YakFull8300

Coding Is solved, Bugs are Not Yet Solved

A post responding to a claim that coding is solved but debugging is not, with the original poster noting that even hand-written code had bugs. The community largely dismissed the framing as out of touch, with commenters comparing it to saying cruise control solves driving. The discussion touches on how AI has made code generation accessible but software engineering — the broader discipline of building reliable, maintainable systems — remains unsolved.

Interesting Points
  • The original claim that coding is solved but bugs are not was widely mocked as a category error, with commenters noting that hand-written code always had bugs too.
  • One commenter compared the claim to saying cruise control solves driving — useful but nowhere near solving the broader discipline.
  • The post was contextualized as a reply to a discussion about Claude Code not hot-reloading an agent designed to run 24/7.
Top Comments

Oh, coding is solved but just not the most notoriously difficult part. Cool, thanks

u/DeterminedThrowaway (174 points · permalink)

I can do brain surgery, the people just don’t survive.

u/crustyeng (53 points · permalink)

Coding is mostly solved. Software engineering isn't

Taking your first example, it is more like saying that cruise control is solved. It's very useful to driving, but not anywhere near "solving" driving

u/daniel-sousa-me (20 points · permalink)


I gave Qwen 3.8 27B a reverse-engineering job I assumed needed a frontier model, and it finished in 30 minutes

383 points · 35 comments · r/singularity · by u/yogthos

Screenshot of reverse engineering results

A user tested Qwen 3.8 27B on a complex reverse-engineering task they assumed would require a frontier model, and the local model completed it in 30 minutes. The community discussed how reverse engineering is an ideal use case for LLMs because the binary itself provides a complete specification for the AI's output, making both the goal and validation criteria well-defined.

Interesting Points
  • One commenter described a full workflow using Cline + Qwen + ghidra-headless-mcp for static analysis and Arkana for dynamic analysis, with Qwen cracking software without difficulty.
  • The community noted that RE is a particularly strong use case because the thing being reverse-engineered provides the complete spec for the AI's output.
  • Users reported that Qwen 3.8 27B fixed tool-use regression issues that affected the previous 3.6 version, which forgot how to use tools when only 1/4 through its context.
Top Comments

RE is a fantastic use case in my experience, because the thing you're REing itself provides the complete spec for the AI's output

Means that both the AI's goal, and how to validate it, are well defined out of the box

u/jesusrambo (80 points · permalink)

It's very good at re, I've been using an abliterated version in concert with deepseek-v4-pro and headless ghidra and it cracks software without breaking a sweat.

u/karlnuw (58 points · permalink)

My full workflow is Cline + qwen + ghidra-headless-mcp for static analysis and Arkana for dynamic analysis, and this as a Cline rule. I had Codex set everything up for me in one folder, I throw the binary in there and Qwen gets to work. If it's struggling because it's a massive binary I switch it out for deepseek.

u/karlnuw (49 points · permalink)

I saw this headline and it inspired me. I've been meaning to RE the old 8088 version of Elite. Unlimited free credits on 0x Alpha this weekend, so I just set it to max and had it crunch for half a day to produce a fully-documented, fully-labelled assembler output that recompiles to a byte-exact binary.

u/QING-CHARLES (6 points · permalink)

I ran low on anthropic usage this weekend in the middle of a bunch of work well-specced by Fable and switched Claude Code to qwen 3.8 max $68/mo coder subscription.

At one point last night I had 5 sessions working on the pre-designed work and some data migrations. They were noticeably slower, but the quality of work didn't suffer enough to notice. I'm going to keep at least the basic subscription for fallback.

u/tribat (5 points · permalink)


Robot plays ping pong with Ding Ning (2016 Olympic champion)

344 points · 71 comments · r/singularity · by u/averagebear_003

Video thumbnail of a robot playing ping pong with Olympic champion Ding Ning

A video shows a humanoid robot playing a competitive game of table tennis against Ding Ning, the 2016 Olympic champion. The demonstration showcases the robot's real-time visual processing, motor control, and adaptive decision-making at speeds that challenge a world-class human athlete.

Interesting Points
  • The robot is playing against a 2016 Olympic gold medalist, one of the greatest table tennis players in history.
  • The demonstration requires real-time visual tracking of a ball moving at speeds exceeding 100 km/h, combined with precise motor control for paddle positioning and swing timing.
Top Comments

Ding Ning seems to be holding back

u/whoknowsifimjoking (1 points · permalink)

I've seen enough. Give him a weapon

u/Individual_Review771 (1 points · permalink)


Amazon kept shutting down my tablet, so I spent $266 on four AI models to own it

322 points · 25 comments · r/singularity · by u/yogthos

A user spent $266 on four different AI models to regain full ownership of their Amazon Fire tablet after the company repeatedly shut it down remotely. The post sparked discussion about the tension between open-source AI empowering individual users versus corporate control, with many noting that devices should come with root access by default since the buyer owns the hardware.

Interesting Points
  • The user leveraged multiple open-source models to bypass Amazon's remote shutdown mechanism, demonstrating the practical power of accessible AI tools.
  • Commenters highlighted the broader concern that Amazon's software can shut down hardware the owner had bought, while the owner needs root to stop it.
  • The discussion touched on how open-source LLMs have the capacity to significantly disrupt the software industry and cause major shifts in market dynamics.
Top Comments

Amazon's software being able to shut down hardware the owner had bought, while the owner needs root to stop it, is more concerning

u/mridugup20 (115 points · permalink)

This is great and I think represents the fear of open source and let er rip approach. Empowering the population comes at then expense of corporate power. The trade off (huge increases in innovation, economic benefit ect) is enticing, but expect a cat amd mouse game to continue.

u/Comfortable_Car6562 (88 points · permalink)

Principle

u/Efont93 (35 points · permalink)

Insane work. Even insaner is, that devices don't come with root by default, after all I am buying the device.

u/Technical-Earth-3254 (17 points · permalink)

This is just a theory, but I'm guessing Amazon doesn't make any money selling the tablets. They make money on the subscriptions and walled-garden aspect, so locking you out of your own device is super-important to their business model. I don't know of many companies selling tablets besides Amazon, Apple, and Samsung. I wonder if I could buy a brand new, root-accessible tablet today. My guess is, if I can, it would cost 3 times as much as the same spec tablet from those 3 manufacturers, because they won't be able to make up the profits on the back end with subscriptions.

u/chuckaholic (8 points · permalink)


Sam Altman with some sad statements about AI

310 points · 198 comments · r/singularity · by u/SwingDingeling

Sam Altman admits he was wrong about the speed of AI adoption, saying he expected rapid disruption to software businesses after GPT-4 in 2023 but found the economy has significant inertia. He notes people keep doing the same things and using the same tools, which he frames as a positive that will make the transition smoother and slower. He concludes that everyone has been too ambitious on timelines.

Interesting Points
  • Altman specifically said he thought 'very quickly after' GPT-4 there would be 'much more disruption, software businesses up for grabs right away, than it turned out to be'
  • He characterizes the slow adoption as 'actually a positive in many ways' that will 'make this big transition go smoother and slower'
  • Commenters note this means bleeding-edge adopters will pull further ahead, though one points out you need to do something monetarily productive with new technology, not just use it casually
Top Comments

It also means that people at the bleeding edge of usage will pull further ahead.

u/SawToothKernel (235 points · permalink)

People would pay more attention to AI once it actually changes something in their lives directly. They hear people yelling exponential but they don't see it.

u/DeviceCertain7226 (70 points · permalink)

'This is too fast and dangerous for society!'

'Society isn't adapting like I BS'd about so it's fine'

u/lazyhustlermusic (47 points · permalink)

A source wouldn't go amiss.

Yeah, he's not wrong. I mean, I'm one of the people with aggressive timelines who believes full RSI and AGI-level systems are right around the corner (next year.) But one can see clearly that even with those, most people just won't be affected. Eventually the results of those systems will filter down into the medical field and others, and that's when people will be affected. That's assuming we don't get fully autonomous ASI, which would be a different thing altogether.

u/lovesdogsguy (33 points · permalink)

I mean I hear what he's saying. Inclined to agree.

But he is a corporate figurehead. You gotta take everything he says with a grain of salt.

u/bear_Prune8771 (30 points · permalink)


I fine tuned Gemma 4 12B for a 2.7x improvement on tool calling because I can't fit anything else comfortably into my 16 GBs of Vram

243 points · 34 comments · r/LocalLLaMA · by u/TheOneWhoWil

I fine tuned Gemma 4 12B for a 2.7x improvement on tool calling because I can't fit anything else comfortably into my 16 GBs of Vram

A user with only 16GB of VRAM fine-tuned Gemma 4 12B using QLoRA SFT to achieve a 2.7x improvement in tool calling performance. The training used glaiveai/glaive-function-calling-v2 combined with zai-org/AgentInstruct datasets, yielding over 5,000 formatted examples. The model's loss dropped from 1.695 to 0.120 over 3 epochs at rank 16 with a batch size of 32. The post has generated interest from the community for sharing practical fine-tuning details that others can replicate.

Interesting Points
  • The training used glaiveai/glaive-function-calling-v2 with all splits of zai-org/AgentInstruct, yielding over 5,000 examples that had to be manually formatted.
  • QLoRA SFT was applied at rank 16 over 3 epochs with a batch size of 32, reducing loss from 1.695 to 0.120.
  • The poster noted the main effort was cleaning and formatting datasets for Gemma 4's template and converting ReAct-style trajectories into tool function calls.
  • Community members asked about the model's ability to make single sequential tool calls versus multiple simultaneous ones.
Top Comments

Oh, go you. You're awesome. Were you able to belt the behavior out of it where it only makes a single tool call at a time?

So many of these new agentic models are able to make three or four sequential tool calls because they know that that's a thing, whereas Gemma never seemed to do that.

u/PossessionUsed7393 (30 points · permalink)

Very cool. Care to share details like datasets, hyperparams, scripts? SFT plus some type of RL?

u/DinoAmino (7 points · permalink)

Basically it's glaiveai/glaive-function-calling-v2 with all splits of zai-org/AgentInstruct which got me over 5 thousand examples that I had to format

I used QLoRA SFT to train the model at rank 16 over 3 epochs and a batch size of 32. Loss went from 1.695 -> 0.120

I don't have the scripts up on github or anything, probably should've just for the green boxes lmao. I didn't do anything that interesting so you aren't missing out on much. Most of the code was actually just for cleaning and formatting the datasets for Gemma 4's template and turning the ReAct style trajectories into tool function calls and reasoning. If you really want the scripts I could send them to you

u/TheOneWhoWil (14 points · permalink)


Is it really wrong to use AI to post or comment on Reddit?

240 points · 262 comments · r/ChatGPT · by u/AmeniNuretemo

A Japanese Reddit user asks whether using ChatGPT as a translation and writing assistant to participate in English-language discussions is problematic. They emphasize that the opinions are entirely their own and AI is used only to bridge a language barrier. The post has generated a nuanced discussion, with many commenters supporting this use case while acknowledging that resentment stems from AI being used to generate opinions rather than translate them.

Interesting Points
  • The poster uses ChatGPT to understand English posts and translate their Japanese thoughts into natural English, describing it as the first time they can comfortably participate in global discussions.
  • One commenter noted the resentment stems from a different application: AI can only translate, or it can do additional cognitive work like formulating opinions, and others have to guess where on that spectrum a given post falls.
  • A commenter noted the poster was asking on a pro-AI subreddit and suggested trying r/antiai for a different perspective.
Top Comments

I think this is one of the better use cases of AI personally

u/CorpulentPutrescence (299 points · permalink)

The opinions are still mine. I'm not asking AI to decide what I think. I mostly use it as a translator and writing assistant.

But I've noticed that some people react negatively as soon as they think a post was written with AI.

So I'm genuinely curious: do you think using AI this way is a problem?

You describe a very specific, disciplined use-case. I agree: It's fine in this case.

The resentment stems from a different application. AI can only translate for you, or it can do additional cognitive work, like formulating your opinion. Only you know where on that spectrum you land. The others have to guess.

u/-Spzi- (116 points · permalink)

If translation isn't an acceptable use case for language models, then I guess nothing is

u/traumfisch (92 points · permalink)

For me, this is a very good use of AI. With pre-GPT translators, you could never be sure if it translated something correctly. (You still can't be 100% sure, but ~97% of the time it works).

However - you're asking this question on a pro-AI subreddit. Try asking the circlejerk that is r/antiai , and see how/if the replies differ. Maybe they'll tell you that you should hire a human translator for $100/h.

u/Elektrycerz (51 points · permalink)

Cancer could be 100% cured by ai and reddit will discredit it because its done by ai.

At this point its a cult of ignorance by choice

u/Such--Balance (41 points · permalink)


54 more Reddit stories

Updates: 05:30 AM PDT · 08:30 AM PDT · 11:30 AM PDT · 02:30 PM PDT · 05:30 PM PDT