| ▲ | taylorfinley 2 hours ago |
| Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time. Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3... |
|
| ▲ | spankalee 2 hours ago | parent | next [-] |
| 3.8 Flash is just quite good, and so is the Antigravity harness. I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible. |
| |
| ▲ | mapontosevenths 2 hours ago | parent | next [-] | | Even if agy was the best (it's not, and is missing basic features) you wouldn't rather have a choice? I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around. | | |
| ▲ | drusepth 2 hours ago | parent | next [-] | | What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me. I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again. | | |
| ▲ | walthamstow 2 hours ago | parent | next [-] | | I have used it for little more than 6 hours or so in total but I'm pretty sure it doesn't have compaction? | | |
| ▲ | KeplerBoy 2 hours ago | parent | next [-] | | How else would it work? Less technical people don't even watch their context usage. | | |
| ▲ | macNchz 38 minutes ago | parent [-] | | In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model's window. That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff. | | |
| ▲ | Vacyyyy 15 minutes ago | parent [-] | | Do you have experience with OAI's, it's been known to be to be good for a while now, going off public consensus and my experience. |
|
| |
| ▲ | honr an hour ago | parent | prev [-] | | It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation. |
| |
| ▲ | sarjann 2 hours ago | parent | prev | next [-] | | Auto mode? | | |
| ▲ | KeplerBoy 2 hours ago | parent [-] | | It absolutely has auto mode. | | |
| ▲ | smartbit an hour ago | parent | next [-] | | agy cli does not have auto mode. I've tried and tried and tried to work with agy cli sandbox-mode and just failed. agy --dangerously-skip-permissions
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli. IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate. | |
| ▲ | SomaticPirate an hour ago | parent | prev | next [-] | | I think it only has "--dangerously-skip-permissions"
Claude and codex auto mode will reject certain actions.
No secondary check on gemini/agy AFAIK | |
| ▲ | levelZero an hour ago | parent | prev | next [-] | | Via cli switch, but in process w/o fine graining? If so please tell | |
| ▲ | LoganDark an hour ago | parent | prev [-] | | Auto mode means that another model reviews tool calls to attempt to disallow less safe ones. It's different from bypass permissions mode which typically just doesn't filter at all. |
|
| |
| ▲ | arizen 2 hours ago | parent | prev | next [-] | | Does it have /goal feature similar to Codex? | | | |
| ▲ | esafak 2 hours ago | parent | prev [-] | | I use a variety of models for various subagents. I don't want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better. |
| |
| ▲ | eloisant an hour ago | parent | prev [-] | | There is a pi plugin to use agy directly from it. | | |
| ▲ | mapontosevenths 41 minutes ago | parent [-] | | You get banned if they catch you. | | |
| ▲ | tcoff91 27 minutes ago | parent [-] | | Yes and I've seen reports of it being an ENTIRE GOOGLE ACCOUNT BAN. I don't want to mess with antigravity because my google account is too entrenched in my life. |
|
|
| |
| ▲ | moecables 26 minutes ago | parent | prev [-] | | I use Antigravity but for some reason, `agy` in the command line feels very bad/incapable of doing things. I can't quite explain it but the most common issue I run into it is just hanging on being unable to finish a tool call |
|
|
| ▲ | IndeanCondor 2 hours ago | parent | prev | next [-] |
| Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight. |
| |
| ▲ | seanthemon an hour ago | parent [-] | | Gemini for day-to-day and top-of-head queries and claude for the real beefy work |
|
|
| ▲ | yegle 2 hours ago | parent | prev | next [-] |
| For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly. And I saw it do this twice, once for Android 14 and once for Android 16. I think this is just within 3.8 flash's capabilities. |
| |
| ▲ | p_l an hour ago | parent [-] | | 3.8 Flash (but also last two ones) have really strong preference for dissecting binaries with quick thrown-together bits of python in my experience. Including going first for decompiling AGY binary instead of searching the web for documentation... | | |
| ▲ | IshKebab an hour ago | parent [-] | | Astra also really loves reverse engineering binaries. I guess it's one of those things that isn't that complicated but is super tedious, and tedium means nothing to AI. |
|
|
|
| ▲ | amanguliani 2 hours ago | parent | prev | next [-] |
| Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me |
| |
| ▲ | onlyrealcuzzo an hour ago | parent [-] | | Sol 6.1 is quite good, but damn is it slow. I'm using it to run overnight tasks, and that's it until my quota runs out. Canceled my subscription. | | |
| ▲ | amanguliani 25 minutes ago | parent [-] | | WHY ARE ALL OPENAI MODELS SO CHATTY - i thought claude kept going on, then i literally put it in claude.md that summarize your thinking in 200 words or less and tell me in points what you did and what's next. Did the same for CODEX - nope still keeps effing going on and on and on |
|
|
|
| ▲ | gottorf 2 hours ago | parent | prev | next [-] |
| My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics. |
| |
| ▲ | staticman2 an hour ago | parent | next [-] | | The web version of Gemini is awful at search but I don't think that's the models fault. | |
| ▲ | MILP an hour ago | parent | prev | next [-] | | I'm also not using it for coding but I've found Flash 3.8 to generate much better HTML output than Sonnet or Opus. | | |
| ▲ | robobo96 5 minutes ago | parent [-] | | Only html or also css? Opus seems a bit more creative than most other models i've seen. |
| |
| ▲ | mattjoyce an hour ago | parent | prev [-] | | Hallucination seems a very dated term. | | |
| ▲ | nkozyra an hour ago | parent [-] | | Why? It's the same concept and root cause it was when we first started using it. |
|
|
|
| ▲ | illwrks 20 minutes ago | parent | prev | next [-] |
| I've been tinkering with Gemini for several months and I think it's great. The most complex things I've had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc. |
|
| ▲ | mapontosevenths 2 hours ago | parent | prev | next [-] |
| Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market. |
| |
| ▲ | ody4242 an hour ago | parent [-] | | what is so terrible with their harness? I've been using gemini cli, now use agy, Pi agent harness, and agent (cursor), and my only real issue with agy was the permission handling, but other than that, it was ok. |
|
|
| ▲ | alightsoul 2 hours ago | parent | prev | next [-] |
| Please tell me you published your findings even as an issue on the llama.cpp GitHub |
| |
| ▲ | taylorfinley an hour ago | parent | next [-] | | Here they are: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3... | |
| ▲ | dominotw an hour ago | parent | prev | next [-] | | he is still closing his jaw | |
| ▲ | otabdeveloper4 2 hours ago | parent | prev | next [-] | | Spoiler alert: the problem didn't actually get fixed despite the jaw on the floor. | |
| ▲ | warkdarrior 2 hours ago | parent | prev [-] | | Why? Anyone can run that prompt. | | |
| ▲ | aspect0545 2 hours ago | parent | next [-] | | Not everybody has access to AI. More than that, every prompt uses insane amounts of natural resources. So why not share it. | | |
| ▲ | FranzFerdiNaN 2 hours ago | parent [-] | | The resources per prompt aren’t that much . Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources. | | |
| ▲ | articulatepang an hour ago | parent | next [-] | | All your examples are private goods: excludable and rival. If one person uses a unit, that prevents others from using them. Patches to open source software are public goods. Your using them doesn’t prevent others from using them. So if you spend resources creating a public good, it’s in everyone’s interest to share it. | |
| ▲ | qmr an hour ago | parent | prev | next [-] | | Yet you participate in a society. | | |
| ▲ | scarmig an hour ago | parent [-] | | If action X takes a million times more resources than action Y, it's silly to focus on or highlight action Y. Seriously: if you are a regular meat eater, your choices use several orders of magnitude more water than even a heavy LLM user. A quip from a comic doesn't somehow erase that or make it irrelevant. |
| |
| ▲ | hexfish 2 hours ago | parent | prev [-] | | Checkmate. /s |
|
| |
| ▲ | luckydata 2 hours ago | parent | prev [-] | | why reinvent the wheel and spend tokens for a problem that has already been solved? |
|
|
|
| ▲ | bel8 2 hours ago | parent | prev [-] |
| I had a similar but less impressive experience recently with Muse Spark 1.3. Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine. It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size. |
| |
| ▲ | seanthemon an hour ago | parent [-] | | Godot encryption is laughably easy to break, there's tons of packages available for it. It's a well known drawback of using godot |
|