Remix.run Logo
Show HN: Ante, a coding agent in a single binary that runs offline(github.com)
101 points by ubermon 4 hours ago | 65 comments
NitpickLawyer 3 hours ago | parent | next [-]

Linking to a github repo for a binary release (no source code related to the agent that I could see) is a bit iffy IMO. You should clarify your intentions or link to something else. Might confuse folks.

OleksandrC 3 hours ago | parent | next [-]

I will leave this here: https://usehax.dev/ GitHub repo: https://github.com/OleksandrChekhovskyi/hax

This is a coding agent implementation I am working on, which delivers what this promises (at least on the "lean" part), except it's actually fully open source, and even more lean (few MBs of runtime memory usage).

MIT-licensed, written in C, multi-provider / multi-model, minimalist approach to system prompt and tools (think kinda like pi, but with a bit more "batteries included", like subagents and background tasks out of the box), polished presentation, inspectable (usable transcript view), etc.

ubermon 3 hours ago | parent [-]

nice, good to see more contributor in this space

lrvick an hour ago | parent | prev | next [-]

I was kind of excited for this until the binary blob. You want me to give your agent binary god access to my computer, and I am not even permitted to see the source code or use my own supply chain security hardened rust compiler stack? What a joke. Hard pass.

ubermon an hour ago | parent | next [-]

we are migrating to public repo. will first see how to ship a binary with source code attached. the release pipeline seem to use the preview repo for attaching source.

queisoy an hour ago | parent | prev [-]

How did the repository get this many stars, and upvoted this much on HN? Boosted by moderators?

ubermon a minute ago | parent [-]

I wonder my self, was notified by a friend, but I am grateful.. Probably because of the Meta's open model release?

ubermon 3 hours ago | parent | prev | next [-]

we put it in the repo README, will add migrate more into public repo as soon as possible.

lrvick 41 minutes ago | parent [-]

It has to be 100% open source public code or we will have no way of proving this is not malware, or secretly swapped out for malware later when your CI/CD system or laptop is compromised. Supply chain attacks happen all the time and with closed code no one will be equipped to spot it when it happens.

Also, security aside, engineers want the freedom to modify and experiment with the tools we rely on.

Tools like this are too important to be closed. Do you want to be Internet Explorer or Firefox?

ubermon 21 minutes ago | parent [-]

Oh no, definitely not Internet Explorer. Will figure out a way to ship with source first to address the security and concern and then figure out how to do the open source development in agentic era later, I was too carried away by the complexity of the latter and ignorant to the former.

lrvick 10 minutes ago | parent [-]

When it is 100% freely licensed open source software I can modify and compile myself, I will be happy to give it a shot.

adastra22 3 hours ago | parent | prev [-]

Linking to a binary is iffy from a security perspective. Linking to a GitHub repository is exactly what HN should do.

messh an hour ago | parent | prev | next [-]

I understand that claude-code takes a lot of memory and that's bad. However, harneses are simple loops, in theory should take very little memory even if written in python or typescript. See for e.g. pi agent

ubermon 14 minutes ago | parent [-]

I love pi and share many vision and value with it. But my view on harness is that it is to capturing the structural mechanism with llm interacting the world. They are a dynamic duo evolving together. The technical depth will continue to grow (e.g. /goal being the new primitive, multi-agent collaboration is basic need)

Yes the core part is a simple loop, but we have all built toy compilers, inference engine, browsers (it is just a curl command eth) etc. The core algorithm is supposed to be simple.

Had a discussion recently: https://x.com/NoCommas/status/2086568454434537710?s=20

thih9 44 minutes ago | parent | prev | next [-]

> We care about the harness, not the model or the prompts.

I wonder if this is a viable approach; after all frontier model providers are betting on the opposite.

Then again, they bundle their harness and offer subsidiary pricing - so maybe they themselves aren’t sure if models are as important.

ubermon 11 minutes ago | parent [-]

my view on harness is that it is to capturing the structural mechanism with llm interacting the world. They are a dynamic duo evolving together. The technical depth will continue to grow (e.g. /goal being the new primitive, multi-agent collaboration is basic need) So it is here to stay. And it is just our focus as we don't have enough resource (yet) to improve the model and I think prompts belong to the user.

Had a discussion recently: https://x.com/NoCommas/status/2086568454434537710?s=20

swrrt 4 hours ago | parent | prev | next [-]

How good is it to work on building games, compared to existing agents? I am building my own game?

andai 2 hours ago | parent | next [-]

They are a bit weird with game development at the moment.

They can one shot entire games, with relatively minor issues.

And obviously asking for small code snippets and integrating them yourself has been well supported for five years.

But in Agent mode... not so much. I was asking frontier models to make simple changes to my Pong game (you know like the one from 1972) and it constantly failed to make simple changes or would break something else in the process.

The main issue is that they can't see what they're doing. Actually one of the agents tried playing the pong game by screenshotting every frame, and it ran for about 20 minutes before I realized what it was doing, and told it to calm down.

It takes about 10 seconds to process an image, so it was running the game at 0.1 frames per second... 600x slower than realtime. The technology is not quite there yet.

If your game is something turn-based though, with discrete States and well-defined transitions between them, they can help out a lot more with that.

ricardobeat 2 hours ago | parent [-]

Depends a lot on what model you use. Claude Sonnet/Opus, Deepseek, Mimo, M3 are smart enough to figure out how to create debug views, add single-frame screenshot and testing harnesses.

1bpp 3 hours ago | parent | prev | next [-]

You aren't building your own game if you have a chatbot do it for you.

marssaxman 3 hours ago | parent | next [-]

A classic tale from the music production world comes to mind:

"I thought using loops was cheating, so I programmed my own using samples. I then thought using samples was cheating, so I recorded real drums. I then thought that programming it was cheating, so I learned to play drums for real. I then thought using bought drums was cheating, so I learned to make my own. I then thought using premade skins was cheating, so I killed a goat and skinned it. I then thought that that was cheating too, so I grew my own goat from a baby goat. I also think that is cheating, but I’m not sure where to go from here. I haven’t made any music lately, what with the goat farming and all."

throwlifeaway 3 hours ago | parent | next [-]

Does prompting an AI music generation service count as making your own songs? I noticed you left that one out of your parable.

jermaustin1 2 hours ago | parent | next [-]

As a songwriter, drum hitter, and acoustic strummer, I am not good at drums (can play on beat, though), guitar (can strum all the cowboy chords, though), or singing (no caveat on this one).

So sometimes when I want to hear what my song COULD be if I were good at every one of those, I will use Suno. But I do not let it change my lyrics, or chose it's own arrangement. I will give it a VERY ROUGH demo (out of key singing, basic drums and strums, and it will "polish" my turds into really shiny turds.

marssaxman 2 hours ago | parent | prev | next [-]

Does snapping a photo with a camera count as making your own painting?

Seich 2 hours ago | parent [-]

We don’t call the end result a painting though, we call it a photo. Theres probably some value to distinguishing between them.

andai 2 hours ago | parent [-]

A set of pleasing qualia.

bitpush 3 hours ago | parent | prev [-]

If figma draws you a gradient, are you really a designer?

nkrisc 25 minutes ago | parent | next [-]

A designer, yes. A painter, no.

1bpp 2 hours ago | parent | prev [-]

Probably yes; If Figma is deciding for you where the gradient goes, what colors it is, and what its purpose in the design is, then no.

8note an hour ago | parent [-]

it kinda is by showing you tools in a certain order and what defaults and so on

1bpp 2 hours ago | parent | prev | next [-]

No, it's using FL Studio's random sequence generator with a general MIDI instrument, exporting it, then saying it's a ballad you just composed. Or paying a ghostwriter to write on a topic, then saying it's a novel you just wrote.

dionian 2 hours ago | parent | prev [-]

"Good artists copy, great artists steal." - Picasso

stronglikedan 3 hours ago | parent | prev | next [-]

That's like saying you aren't building a house if you use a hammer to drive nails instead of your hand. AI is just a tool like any other, and you use it to build things like you would any other tool.

bigfishrunning 3 hours ago | parent | next [-]

I'd say it's closer to a Roomba then a hammer -- if you just let a Roomba run around your living room, can you say you vacuumed?

simlevesque 3 hours ago | parent [-]

Well, you tell your Roomba "clean the floor" but you don't ask an AI "make a game". You give it very specific instructions.

finghin 3 hours ago | parent [-]

Specific compared to writing procedural code> Barely even by analogy, IMO

One could even argue what defines AI instructability is heuristics as opposed to specifics

ThrowawayR2 3 hours ago | parent | prev | next [-]

A far more accurate analogy would be people saying someone didn't write a book if they hired a ghostwriter, which is generally acknowledged to be true. LLMs are much closer to being a ghostwriter than they are to an inert, non-powered hand tool.

schnevets 3 hours ago | parent | prev | next [-]

I wouldn't even compare it to a tool. People frequently "build their own house" where they sub out 75% of the skilled labor and act as a General Contractor/glorified gopher.

And you won't get purity tests from the layperson: in the end, you're responsible for the build quality so if you tirelessly labor/oversee those teams you're considered capable; if it ends sub-par, then you're a stooge.

throw_m239339 3 hours ago | parent | prev [-]

> That's like saying you aren't building a house if you use a hammer to drive nails instead of your hand. AI is just a tool like any other, and you use it to build things like you would any other tool.

A coding agent is more like a carpenter, a mason, an electrician,... rather than a hammer in that case.

derefr 2 hours ago | parent | prev | next [-]

"A game" is a different abstraction layer from "a piece of software", though. A game has designed mechanics, a scenario (level design, etc.), art/music assets, writing, and so on. I would say that if you're making all of those things yourself, but you're having an AI write the software that executes the game, then you're still "building a game" per se. "Developing a game" even.

Compare/contrast: people who develop games on top of high-level genre-specific game engines like RPG Maker are still considered to be "building a game." What's the difference between using a pre-made purpose-fit engine like RPG Maker, vs. asking an AI (or, for that matter, a contracted software company) to build you a custom purpose-fit engine?

jatora 3 hours ago | parent | prev | next [-]

This is such a horrible cringe and bad faith take.

andai 2 hours ago | parent | prev | next [-]

Eating food doesn't count if you didn't grow it yourself.

perching_aix 3 hours ago | parent | prev [-]

Same idea that sends people down needless game engine development / procrastination rabbit holes.

ubermon 4 hours ago | parent | prev [-]

we have a detailed launch thread explaining and show case exactly this! https://x.com/NoCommas/status/2086835536598351955

jhgik798 3 hours ago | parent | prev | next [-]

no source code

ubermon 3 hours ago | parent [-]

partially, we indent to progressively add source code component to it. the crates/ is sync from the active repo in realtime.

_pdp_ 3 hours ago | parent | prev | next [-]

Considering that ripgrep, git, and, you know, other dev tools are part of the toolbox, then why ship them inside this executable? And, furthermore, if you ship them, then why stop there?

ubermon 3 hours ago | parent [-]

so it is more for being self contained and works out of box if being deployed in a bare linux environment. we tried shell out to `rg` it didn't work very well and instead spending time handling the args parsing and jugging string output, we decided to spend time on building Grep natively for agent. It is just a start~

as for why stop here yep, the goal is to be able to find the best sweet spot in being self contained v.s. all-in-one bloat ware. for example, we still use a bundled `tmux` skill for the orchestration.

plainviewinstru 3 hours ago | parent | prev | next [-]

this would go very hard with a lightweight gui

ubermon 3 hours ago | parent [-]

yes, the goal is to perfect the `ante serve` so it is easy to build gui. We are building one internally to test the protocol version

ubermon 4 hours ago | parent | prev [-]

Hi HN, I'm Mohan from Antigma Labs. Ante is a coding agent that ships as one self-contained ~15MB binary: the TUI, an embedded ripgrep, local PDF/OCR, and a natively managed llama.cpp engine are all inside. No runtime dependencies, no node_modules, no account.

- Ante installs a pinned, checksum-verified official llama.cpp build matched to your machine (Metal on Apple silicon; CUDA, Vulkan, or CPU on Linux) and handles upgrades when the pin changes. - It discovers GGUF files already on disk (~/.ante/models, the llama.cpp and Hugging Face caches), attaches to llama servers already running on local ports, and estimates RAM/VRAM from model size and context window before anything loads. - `ante --offline-model /path/to/model.gguf "prompt"` boots the server, runs the session, and shuts it down. `/offline-mode` does the same interactively; `ante serve --offline-model` loads a model once for many clients. - No API key, no account. Once the model is on disk, inference needs no network at all; set ANTE_TELEMETRY=off and no telemetry is exported either.

On capability, we'd rather publish the number than oversell: we benchmark local models with the same harness and auditable runs as frontier ones, and Qwen3.6 27B (a 17 GB download) scores 56.2% on Terminal-Bench 2.1 across 445 trials (live results: https://antigma.ai/eval). That's a real gap from frontier models. The design bet is that you mix: hosted providers and local live in the same catalog, `/providers` switches mid-session, so sensitive repos or high-volume work go local and hard problems go frontier.

Hosted models work with your own keys or subscription. But nothing about trying Ante requires signing up for anything: download the binary, point it at a GGUF.

Offline mode is under active development and has rough edges with the overview at https://ante.run/local/overview. I'll be in the comments.

felooboolooomba 2 hours ago | parent | next [-]

Here you say:

  > Ante installs a pinned, checksum-verified official llama.cpp
But in README:

  > Ante ships its own inference engine
May I suggest you use the first phrasing in both places. I took it as Ante devs had written their own engine and I doubt I'm the only one.
ubermon an hour ago | parent | next [-]

there is one toy version https://github.com/AntigmaLabs/nanochat-rs README updated.

ubermon an hour ago | parent | prev [-]

the public repo README is updated.

nazgulsenpai 4 hours ago | parent | prev | next [-]

Not sure why this was dead but I vouched. It would be nice if telemetry was opt-in, otherwise this looks awesome and can't wait to try it!

ubermon 3 hours ago | parent [-]

will add those soon!

adastra22 3 hours ago | parent | prev | next [-]

Where is the source code?

ubermon 3 hours ago | parent [-]

for now only some of core crates is migrated, will do so progressively

wwww23 2 hours ago | parent | prev | next [-]

Many people have slow computers, but for agent it is no problem. Only run LLM slow too.

niutech 3 hours ago | parent | prev | next [-]

Why not make it Actually Portable Executable using Cosmopolitan Libc, like Llamafile, to make it run on Windows/Linux/MacOS? Why don't you support Windows with CUDA?

ubermon 2 hours ago | parent [-]

even with power of AI, we are mere human and still slow. Adding this to backlog.

majorchord 3 hours ago | parent | prev [-]

Opt-out telemetry is a hard no for me, sorry.

nextblock 3 hours ago | parent | next [-]

Agreed! When a tool is explicitely marketed for offline use, opt-out telemetry feels especially contradictory. Should Definitely be opt-in by default...

stronglikedan 3 hours ago | parent | prev | next [-]

Right? How hard is it to just ask a single opt-in question during installation. Opt-out just seems lazy, especially if I have to dig though configs to get to it.

ubermon 3 hours ago | parent | prev [-]

feedback received, it was carry over from the preview dev build.