Remix.run Logo
What's the best programming language for coding agents?(danluu.com)
71 points by chaychoong 10 hours ago | 57 comments

Related: Which programming languages are most token-efficient? - https://news.ycombinator.com/item?id=46582728 - Jan 2026 (91 comments)

michaelteter 26 minutes ago | parent | next [-]

I'm not sure I trust a source that says "just 70 tokens average, nearly half of Clojure (109 tokens)".

There's no reason to add the phrase "nearly half of", and there's especially no reason to add it when it's significantly far away from half.

But on the main topic, I still feel that Go is an excellent choice for LLMs. There is pretty much just one way of doing most things, and the available training data is pretty consistent. This is very different from Python, where training data is polluted (I presume) with tons of code written by non-software engineers and demonstrating many different ways of doing the same thing.

Also a big plus for Go is the tooling. Fast compiles and good linting shortens the iteration cycle time, resulting in less need for me to tell the LLM to correct mistakes.

For some reason, most LLMs I've used default to wanting to write Python. I have to repeatedly teach them to use Go unless there is a very compelling reason to choose otherwise.

I would personally rather see and use Clojure, but I don't feel its ecosystem would provide the same benefits as Go, including obviously the easy single binary distribution.

gr_norm 2 hours ago | parent | prev | next [-]

It's not clear to me how useful of a signal replicating existing pieces of well-known software is for this kind of evaluation, given what we know about how effectively LLMs can retrieve data from their training corpus and style-transfer it across different settings (programming languages here). That would explain their convergence in ability across different languages on the tasks in this post. I'd be far more interested in people's real-world experiences.

lowbloodsugar 28 minutes ago | parent [-]

I tried writing an AI harness in Python. Seemed the obvious way to go. Tons of libraries. Libraries for talking to model APIs. Libraries for context and conversation management. Libraries for talking to MCPs. It is the language for LLMs!

It was a shit show and just couldn't write anything that would not crash. Super confident it had done a good job. Full of random bugs. A UI needs interactivity, interruption, handling exceptions. It produced some of the worst code I've ever seen. And looking at the libraries' code: also some of the worst code I've ever seen.

I switched to rust + tauri. In about three person weeks of work I have UI with forking conversations, tool use with built in grepping, tons of quality tools. It's more productive (for me) than Claude Code (CLI or desktop).

MichaelNolan 2 hours ago | parent | prev | next [-]

Ive been amazed at how well LLMs are at writing Gleam[1] and Lustre[2]. Compared to a mainstream language, there is basically zero gleam code in the training data.

I have no evidence to back this up, but I suspect that languages that are good for humans[3] will be good for LLMs. Compiled, strongly typed, statically typed, immutable, pure functions, pattern matched, memory safe, etc.

[1] https://gleam.run [2] https://lustre.hexdocs.pm [3] Yes I realize that languages features that are "good for humans" is a hotly debated topic. That's just my personal list for what I like in a language.

maleldil an hour ago | parent | next [-]

Gleam has been stable for over two years, so maybe it's been long enough that LLMs have internalised the documentation.

Given it's a language that doesn't really contain any groundbreaking ideas[1] (the closest is 'use' IMO), it's possible LLMs can reuse patterns from other functional language.

[1] This isn't criticism. I love how Gleam turned out.

ojkelly 35 minutes ago | parent | prev | next [-]

I’ve been developing a language for a few years, and even with incomplete semantics and a simple one page example LLMs don’t have much trouble writing it.

I think the language/syntax has an impact, but the tooling around it will be most important for LLMs, in the same way it is for humans.

jdiff an hour ago | parent | prev [-]

That's not a take I was expecting to find here. I've found most LLMs absolutely dreadful when it comes to Gleam, to the point that I most often disable even inline autocomplete when working in Gleam codebases.

Too often I find them getting pulled into larger ruts in the training data and trying to insert language features that don't exist (ifs, loops, and syntactic constructs) from more popular languages like TypeScript and Rust. Do you not experience other languages getting partially substituted in when you have LLMs write Gleam?

MichaelNolan 38 minutes ago | parent [-]

I suspect it depends a lot on the llm/harness being used. But when I use Opus/cc or sol/codex, at the end of the turn everything compiles, passes tests, and passes lint. I never even look at code that can't compile. Maybe the LLM is generating weird stuff in-between, but I don't see it.

What you're describing feels like my experience back in 2024/25. Back then I was using a llm auto complete or the chat interface, and I would get weird stuff all the time. (not just gleam but any language).

nogha an hour ago | parent | prev | next [-]

Cool seeing Guards of Atlantis 2 here.

One thing that often happens with board games is rule issues in translations. Specifics that are clear in one language get lost in translation. Wolff Designa is out of Latvia. So not surprised there are some hard to interpret rules.

It’s interesting that LLMs struggle with the board game rules like we do. I think game designers should get the llm to teach them from their rulebook. If an LLM can’t understand the rules good chance people will also be confused.

summarybot 15 hours ago | parent | prev | next [-]

Cool line of questioning, but one piece of information is pivotal and critically not-yet-included: equivalent accomplishments in each language. For example, if I want to write standard things: web server, memoized fibonnaci, recipe search engine, what's the length-and-density of these outputs for each language? I think that would add in some ~normalization.

quinnjh 2 hours ago | parent [-]

Strongly agree- this is how I “evaluated” languages pre-agents. though I suspect this would bias results in favor of whatever has best signal to noise for boilerplate from stackoverflow/reddit , rather than what LLM’s “””reason””” best with. (Presuming those aren’t quite one-and-the-same)

dang 3 hours ago | parent | prev | next [-]

Related:

Which programming languages are most token-efficient? - https://news.ycombinator.com/item?id=46582728 - Jan 2026 (91 comments)

genxy 2 hours ago | parent | prev | next [-]

What is the best language for the user of the LLM?

What is the best language to have high quality correctness oracles so that the user doesn't have to babysit the LLM and do lots of manual testing?

frollogaston 2 hours ago | parent [-]

JS is the best tradeoff between succinct and easy to understand. Python is next but has some rough edges that they avoided in JS.

3eb7988a1663 an hour ago | parent [-]

You are going to have to give more support for those assertions. I write Python every day, and never would I call it a good candidate for the clankers. Pretty much any dynamic language would be ruled out, as there is too much implicit logic which makes it harder to understand what is happening.

maleldil an hour ago | parent | next [-]

Python with a strict linter and type checker (eg ruff with the right lints on and ty with its stricter settings, or strict pyright if performance isn't too bad) works very well. Most of Python strengths (concise, large ecosystem, well represented in the LLM training data) while having good static analysis.

frollogaston an hour ago | parent [-]

You don't need that, gets in the way more than it helps. Even Typescript isn't really needed, but at least it's decent devex unlike the Python typing stuff. What really helps is testing.

frollogaston 40 minutes ago | parent | prev [-]

What's better for this, Go? That's the least verbose static one, and it's still a lot more verbose without helping you understand any better what it's doing. It's just faster. That's the real benefit of static types.

nylonstrung 6 hours ago | parent | prev | next [-]

One thing worth noting is that syntactic density doesn't necessarily mean cheaper because because symbols don't chunk/tokenize as well as plain English

What I see from results like this is that the delta between languages is small enough now that it's hard to justify not not using something like Rust for the performance and correctness benefits if you're using LLMs and it fits the domain

clbrmbr 2 hours ago | parent | prev | next [-]

I discovered last week that Fable 5 can write perfect xTensa LX7 assembler code without tools or references. Mind blown.

But, when working on a creative graphics task, the results were best in Lua, middling in integer-only C, and underwhelming in ASM in terms of creative depth.

frollogaston 2 hours ago | parent | prev | next [-]

Any good LLM service (not just coding-focused ones) will write and run ad hoc code without being asked if your prompt involves lots of data. Gemini and Claude tend to pick Python with maybe some SQLite. Some of that must be due to portability alone, but it also means they'll make sure the model and tooling are good at those.

pianopatrick an hour ago | parent | prev | next [-]

I'd like to see the results for Ada on these same measures. On the theory that the Ada type system covers more classes of errors than other languages, and so AI can self correct better.

platinumrad 42 minutes ago | parent [-]

Unfortunately for static type weenies like me (and you, presumably), types don't seem to matter at all, or Python and Javascript wouldn't be on top. There's no reason to believe that Ada's type system is so unique that it alone can help AI self-correct, and Rust, Haskell, ML, Typescript, etc. can't.

DarkContinent 2 hours ago | parent | prev | next [-]

Is there a relationship between how good a programming language is for coding agents and how popular it is among humans? If so, wouldn't Python be the best language for agents, since it's is the most popular (and hence has the most context available for models)?

3eb7988a1663 an hour ago | parent | next [-]

Pick something slightly esoteric (eg Haskell) and the quality of public code is very high, because you only have enthusiasts writing it. Choose something taught in schools (Python) and you are going to find 10,000 traveling salesmen homework problems and Django todo applications.

Not sure how you thread the needle on the quality vs quantity dynamic.

throw-the-towel 2 hours ago | parent | prev | next [-]

As much as I love Python, JavaScript (including TypeScript) is probably more popular.

frollogaston an hour ago | parent [-]

That and JS code is more readily available in the source of tons of webpages, not hidden away in some backend

maleldil an hour ago | parent [-]

Wouldn't most frontend JS in Web page sources be minified?

frollogaston an hour ago | parent [-]

The logic is still there, it's not meant as obfuscation. Also plenty of sites don't minify cause that involves a whole toolchain.

Sha1rholder 2 hours ago | parent | prev [-]

There is definitely a relationship. But I personally believe that once the training corpus reaches a certain scale, the returns exhibit diminishing marginal effects, to the point that multiplying the data volume cannot surpass something essential inherent in language design. (Asked an LLM to help me with the translation, so forgive my expression)

aleph_minus_one 6 hours ago | parent | prev | next [-]

> Dynamically typed languages generally have a lower LLM token cost than traditional statically typed languages because omitting explicit type declarations makes the code more compact.

If this was true, the programming languages that are very much on the left side of

> https://danuker.go.ro/programming-languages.html#non-math-ma...

> https://danuker.go.ro/programming-languages.html#overall-map

should be very ideal for LLMs, in particular if they are dynamically typed.

What I can tell you is: I experimented with AI prompts for generating Wolfram (Mathematica) code using some LLMs, and I can tell you that the results were very disappointing: in my experience LLMs have difficulties with programming languages that are

- very concise, and

- for which there is less code publicly available.

Wolfram (Mathematica) is a good example of such a programming language.

acchow an hour ago | parent | next [-]

> omitting explicit type declarations makes the code more compact.

I guess this ignores languages with type inference? Hindley-Milner and others

JoeyJoJoJr 4 hours ago | parent | prev | next [-]

I’ve actually found Sol delivers great results with Odin, despite there not being much Odin code available. I think it is able to work well with it because:

- It is a rather simple language - It has a lot of very useful libraries already built in.

With just a single main.odin file you can do a heck of a lot stuff, which LLMs seem to like.

ch4s3 2 hours ago | parent | next [-]

It’s interesting I’ve been surprised by how well Claude sonnet can write code in a language I’m developing that probably has no code in the training set. It seems like anything with syntax like python/ruby/elixir is pretty LLM friendly, and layering on a HM type system seems to help catch most errors.

aleph_minus_one 4 hours ago | parent | prev [-]

> I think it is able to work well with it because:

> - It is a rather simple language - It has a lot of very useful libraries already built in.

> With just a single main.odin file you can do a heck of a lot stuff, which LLMs seem to like.

Also Wolfram/Mathematica has an insane amount of useful libraries already built in (there even exists the saying "Python is 'batteries included', Wolfram is 'spaceship included'"), and also there in a single file you can do a heck of a lot stuff.

On the other hand:

- LLMs tend to hallucinate non-existing function when you ask an LLM to code something in Wolfram that is not commonly done (concerning this point, nevertheless keep in mind that Wolfram is often used for "one-of-a-kind programs", i.e. for writing very specialized programs that have possibly never been done before).

- Wolfram code tends to be quite dense.

- If there is a small mistake in Wolfram code, the code typically simply won't work.

petra 2 hours ago | parent [-]

Is there a way in Wolfram to check whether all function names exist ? And than give it as feedback to the llm?

aleph_minus_one 41 minutes ago | parent [-]

> Is there a way in Wolfram to check whether all function names exist ?

There is a way to check whether a symbol has been defined:

  ValueQ[FunctionName, Method -> "SymbolDefinitionsPresent"]
See https://reference.wolfram.com/language/ref/ValueQ.html

Replace FunctionName by the function name that you want to check.

frollogaston 2 hours ago | parent | prev [-]

Training data is a factor too

lowbloodsugar 44 minutes ago | parent | prev | next [-]

First, How fast is the Zstd decoder in python at runtime? If rust and python are essentially the same cost, then chose rust.

Second, I am surprised that python scored slightly better than rust. My own experience is that, when programming python, Claude would spend so much more time dealing with the code not working at runtime, while for any given rust problem, rust would likely fail at compile time, iterating faster and taking less tokens. Some tasks in python it just completely failed at, writing awful garbage. I suspect that is because there is much more awful garbage written in python. (I was trying to write an AI harness. Python seemed like the obvious choice. It was decidedly not).

But in this article, python took slightly less time and tokens than rust for both experiments.

I asked Claude: could you write a decoder, from memory, in python (dont do it, just tell me if you could)

> Honestly: I could write something that's structurally right and would not decode a real .zst file.

> The control flow I'm confident about from memory — frame/block parsing, the literals section dispatch, Huffman weight reconstruction, the backward bitstream reader, the interleaved three-state FSE loop, sequence execution with the repeat-offset rules and the overlapping-copy hazard. I'd expect to get that architecture right, and it would be readable.

So perhaps asking it to do things that are in its memory is not a good benchmark. It was trained with the C "educational decoder, and every third-party port in Rust, Go, Java, JS." and offered a working link [1] to the former.

  [1] https://github.com/facebook/zstd/blob/dev/doc/educational_decoder/zstd_decompress.c
_doctor_love 8 hours ago | parent | prev | next [-]

I love Dan's writing. I really do. But I don't understand why he doesn't have some basic styling on his blog so that it's easier to read.

Kuyawa 10 minutes ago | parent | next [-]

body { margin: 5%; }

That's all it needs, responsive enough for all devices. He can keep his styleless design but margin is always needed.

chiply 5 hours ago | parent | prev | next [-]

I love this take because I had exactly the opposite idea. I thought the combo of remarkably simple text (not even wrapped) with incredible, full width visualizations was chef's kiss. I really like the balance there personally, but I hear you. Does your browser have Reader Mode or something like that? I don't use those tools personally, but I believe they will recast the text parts into something that renders optimally for reading (ideal font size, number of characters per line, etc....).

scared_together 5 hours ago | parent | prev | next [-]

It may be an artistic/engineering choice to demonstrate what minimizing bloat to an extreme degree looks like.

https://danluu.com/web-bloat/

nicebyte 2 hours ago | parent | prev | next [-]

reader mode helps.

9rx 7 hours ago | parent | prev [-]

Users being able to supply their own stylesheet is a core tenant of CSS. Go nuts and make it look however your heart desires!

_doctor_love 6 hours ago | parent [-]

Supply my own stylesheet? No thank you, I'm not here to do work for free.

9rx 6 hours ago | parent [-]

Is doing something for yourself really working for free? That's an interesting take. But I can understand why you don't want this for yourself, so enjoy the page in all its splendour as it is already!

lyall an hour ago | parent | next [-]

> Go to restaurant

> Order food

> Food comes out as raw, unprepared ingredients

> Complain to chef

> Tells me to go cook it myself

> wtf, I'm not here to do work for free

> "Is doing something for yourself really working for free?"

_doctor_love 6 hours ago | parent | prev [-]

So every person who reads Dan's blog and finds the layout too dense, they should write and maintain a stylesheet for his site?

And every person globally should do this as well for any other website that doesn't have a good default reading experience?

dash2 3 hours ago | parent | next [-]

If most readers of danluu don’t find that, then yes!

lemming 3 hours ago | parent | prev [-]

I mean, if it really bothers you you could fairly trivially apply picocss or whatever to it using a user stylesheet. That is so little effort that calling it working for free would be disingenuous to say the least.

tclancy 2 hours ago | parent [-]

Multiple people, me being the third or fourth, are not feeling the default layout and you all read that as a signal it's working as intended?

lemming an hour ago | parent [-]

No, just that it’s easy to change for those that don’t like it. Clearly some people do like it (including Dan, presumably).

cynicalpeace 2 hours ago | parent | prev [-]

I've long suspected that LLMs will just output pure bits eventually

rytill 28 minutes ago | parent | next [-]

Why would this be the case when the text that produces binaries (code) is usually both more token efficient and vastly more effectively organized for modification/extension?

Unless by bits you just mean text in general, or any data since it’s all bits, in which case what you’re saying is trivially already true.

It seems like you’re saying that long term LLMs will output pure machine code as the most effective way to use them.

hankbond 2 hours ago | parent | prev | next [-]

well they can natively converse in base64

nicebyte 2 hours ago | parent | prev [-]

are you implying that text is impure bits?