Remix.run Logo
Protobuf has LSP support. You're welcome(buf.build)
109 points by theanonymousone 7 hours ago | 80 comments
jvolkman an hour ago | parent | next [-]

> Protobuf now has modern IDE support for the first time

Weird post. I built IntelliJ protobuf support [1] while at Google like 10 years ago, and it started shipping by default with IntelliJ in ~2021. Maybe that's not considered "modern".

[1] https://github.com/jvolkman/intellij-protobuf-editor

alecthomas 3 hours ago | parent | prev | next [-]

What an oddly arrogant post, there's been a Protobuf LSP available for years: https://github.com/lasorda/protobuf-language-server

llm_nerd 3 hours ago | parent [-]

The bizarre "you're welcome" in the title is so obnoxious that I was sure it would be related to some sort of funny twist in the blog entry. Nope, they really thought that was appropriate.

colechristensen 10 minutes ago | parent [-]

surely we're not a community so unfamiliar with lacking social skills and this aberration isn't at all important to the topic at hand

lacoolj 3 hours ago | parent | prev | next [-]

"You're welcome" is absolutely hilarious to read from a company post.

bmitc 3 hours ago | parent | next [-]

It's gotten out of hand. Companies making money think they do everyone a favor.

tredre3 3 hours ago | parent | next [-]

They are doing us a favor, there's no denying that. Just like dozen of other providers did them a favor by providing them with the free software pieces they relied on to build their product.

Of course it's poor taste to add the "You're welcome". But poor taste is par for the course for such startup, unfortunately.

imoverclocked 2 hours ago | parent | prev [-]

Despite interesting legal definitions, companies are just groups of people.

I also read the title as having an air of self-importance or maybe passive-aggression. If you read it a little more dispassionately then it's just an awkward way of stating that this is a gift.

Why is this a gift? The company does not need to give it to you, though it arguably benefits from doing so. It did so with some definition of "free."

trial3 3 hours ago | parent | prev [-]

Your HN comment has a reply. You're welcome.

williamcotton 4 hours ago | parent | prev | next [-]

I looked at the dependencies and noticed that it wasn’t using an existing Protobuf parser which means they reimplemented the parser from scratch. Perhaps due to a lack of error recovery in the existing implementations? I don’t have the energy right now for further inspection.

It is definitely best to reuse the parser for the runtime when implementing an LSP but to do so properly means implementing the parser itself as a standalone library. Even better is shipping the semantic analysis as well!

Implementation drift is definitely an issue.

But great project anyways, just wanted to put my thoughts on the matter into the conversation!

vyaa 4 hours ago | parent | next [-]

I think its worth considering but I don’t agree it’s always best to use the language parser.

A language parser should be correct. A LSP parser should be fault tolerant. My understanding is you cant have both.

lalitmaganti 4 hours ago | parent [-]

It absolutely can be done [1] but it is a lot of effort to do high quality error recovery. It's easier in languages with natural "synchronization points" [2], harder in languages which don't have them.

Having built a lot of protobuf tooling, I'd estimate that protobuf largely falls into the former camp; most of your time working with .proto files is operating on fields which terminate using `;` at the end of the line (though multi-line is also possible).

[1] source: I've done it for SQLite SQL at https://github.com/LalitMaganti/syntaqlite/

[2] e.g. SQL naturally has this at statement and expression boundaries which covers almost all of the cases people care about.

afdbcreid 3 hours ago | parent [-]

Adding to this, rustc's parser is both correct and fault-tolerant, yet rust-analyzer ended up building its own for different reasons (needing a CST and not an AST).

williamcotton 2 hours ago | parent [-]

I settled on starting with a CST using rowan, doing semantic analysis, having an agnostic editor services, exposing both an LSP interface with tower, and a WASM interface (for Monaco in the browser [non-LSP!!]) and having the runtime binary doing the interpretation/compilation with the same pathway, with a slight jog, as the editor services. Phew!

adastra22 4 hours ago | parent | prev | next [-]

It is my understanding that basically all LSP implementations use tree-sitter because of its incremental rebuild capability, and this requires re-implementing the parser.

afdbcreid 3 hours ago | parent [-]

That's just not true. rust-analyzer does not use tree-sitter. Neither does gopls. Nor Pylance. Or tsc. In fact I don't know if there is a single popular LSP that uses it (but there probably is).

Many editors use tree-sitter, but that is separate from the LSP.

arccy 4 hours ago | parent | prev [-]

buf is a parser / compiler as an alternative to protoc https://buf.build/docs/migration-guides/migrate-from-protoc/

williamcotton 4 hours ago | parent [-]

Ah, so does the LSP just call out to the Buf binary?

hxtk 3 hours ago | parent | next [-]

The LSP is one of the subcommands provided by the buf binary. My LSP configuration for protobuf is `buf lsp serve`.

williamcotton 2 hours ago | parent [-]

Makes sense, thanks for the clarification.

arccy 3 hours ago | parent | prev [-]

The lsp is a subcommand of buf, from the page: `buf lsp serve`

eterm 4 hours ago | parent | prev | next [-]

Lots of naysayers here, but an advantage of protobuf is that proto files are hand-writeable, and therefore having an LSP for that could be useful.

That said, proto itself dissuades or forbids the kind of common things you might do with a LSP, such as renaming.

Renaming fields is a big no-no [edit: this isn't true, see corrections below], as is doing things like re-ordering fields.

A core idea of proto is that versions are strictly compatible with with previous versions. This itself has limitations and challenges for migrations, but encourages good practice about compatibility that usually gets ignored or hand-waved away in most ecosystems.

I accept however that it's often easy to offload both the re-structuring and the checking of version compatibility to an LLM and let them go at it.

wazzaps 4 hours ago | parent | next [-]

Actually renaming and reordering fields is completely fine, as long as the field ID and type stay the same

ventana 2 hours ago | parent | next [-]

Considered breaking change because JSON serialization will change, and JSON is used quite often.

Reordering fields without changing their ID is indeed a non-breaking change from what I understand, albeit quite a useless one :)

imoverclocked 2 hours ago | parent [-]

> a useless one

If you have a (long) list of fields and you want to keep them lexicographically sorted, being able to reorder is quite useful if you rename a field.

ventana 2 hours ago | parent [-]

Renaming is forbidden though (because JSON and textproto). In Google, it's a documented antipattern to try to make protobuf look "nice" by changing field indices, rearranging fields, etc. — the common ground is that it's better to not do it.

dieortin 4 hours ago | parent | prev | next [-]

Depends, if you use textproto then renaming is problematic too

eterm 4 hours ago | parent | prev [-]

Shit, you're right, I'm trying to remember the one that used to trip us up all the time, maybe it was removing fields?

There was definitely one that kept tripping up the checks and it was something people like to do.

arccy 4 hours ago | parent | prev [-]

renaming is actually pretty fine, if you don't do stuff like json or text encodings. only renumbering fields causes problems.

gjvc an hour ago | parent | prev | next [-]

"You're welcome." is such a smug verbal tic.

unprovable 3 hours ago | parent | prev | next [-]

From protobuf to quantum computing, Google just keeps solving problems nobody has...

ventana 2 hours ago | parent | next [-]

That's not Google. Buf.build is, hilariously enough, a completely separate company that tries to make sense of Google's protobuf, and it's doing a decent job from what I can see.

r_lee 3 hours ago | parent | prev | next [-]

so you're saying protobuf has no purpose?

booi 2 hours ago | parent [-]

i mean...

teaearlgraycold 2 hours ago | parent | prev [-]

Protobuf should at least get credit for enforcing fields be optional.

PufPufPuf 4 hours ago | parent | prev | next [-]

Buf does great work at fixing Protobuf to the point of being just barely usable. A godsend if you're stuck with Protobuf/gRPC on a legacy project.

qweqwe14 3 hours ago | parent [-]

What world are you living in where Protobuf/gRPC are considered "legacy"? ..what?

antonvs an hour ago | parent [-]

In the web dev world, I’ve noticed there are a lot of people who think that JSON is the ultimate modern data format. Typically those same people have barely heard of JSON Schema.

aussieguy1234 2 hours ago | parent | prev | next [-]

I've always thought of protobuf as a Java thing.

It worked great in Java, but not so well in JS, especially with Kafka.

antonvs an hour ago | parent [-]

Protobuf was never particularly a Java thing. You could more accurately say it was a C++ thing. It was initially developed in C++, and used mainly in C++ services at Google. The implementation and early ecosystem were C++-centric.

Later, Go became one of the major Protobuf ecosystems, and today it would be understandable to think it was a Go thing.

gafferongames 5 hours ago | parent | prev | next [-]

While not a direct competitor to protobufs, if you are working in the video game space where struct versioning is not needed, there is an alternative language called "schema" that supports C, C++, C#, Golang, Rust and JavaScript.

https://github.com/mas-bandwidth/schema

satvikpendem 5 hours ago | parent | next [-]

Of all names, they pick schema? That's like calling a programming language "language."

max-privatevoid 5 hours ago | parent [-]

My favorite programming language is called "A Programming Language".

andai 4 hours ago | parent [-]

Reminds me of xkcd tattoo that says in Chinese, "It's what my tattoo says."

jayd16 4 hours ago | parent | prev | next [-]

> video game space where struct versioning is not needed

Save files? Looser than exact version multiplayer?

gafferongames 4 hours ago | parent [-]

Multiplayer games typically deploy both client and server at the same time, and refuse to connect a client if it doesn't speak the exact same protocol as the server.

Thus all the versioning overhead of protobufs is not needed for this wire protocol.

(Yes, games still use versioning everywhere else where it makes sense: save games, asset data, config etc...)

npstr 4 hours ago | parent [-]

Good luck deploying a client to the mobile app stores together with your backend :pain:

But then again, those barely count as games, I guess.

dvtkrlbs 4 hours ago | parent | next [-]

Wouldn’t it make sense to also have the old version of the game live and gradually roll out the new version for mobile

jayd16 2 hours ago | parent [-]

You can do this but if it's too granular (like you have no concept of version compatibility) then it can heavily split your matchmaking.

Plus the headaches of keeping many out of date builds up to date enough to deploy.

Even if you don't care about in game compatibility, all your servers still talk to some centralized data store and that will likely want a single deploy that handles old clients

cyberax 4 hours ago | parent | prev [-]

We do it just fine (not for games). You submit a new binary for review in advance and then do a coordinated release once the review completes.

jayd16 2 hours ago | parent [-]

Then players are locked out of multiplayer until they download a possibily large content patch.

eventualcomp 3 hours ago | parent | prev [-]

Checked the contributor list, please disclose that you're a primary contributor.

ltbarcly3 5 hours ago | parent | prev | next [-]

Watch as I don't use protobuf because it is horrible.

....

Tada!

If you patch clients to google services in Python to use json instead of grpc they get faster and more reliable. A lot faster. Benchmark it!

  def get_json_client() -> CloudLoggingQueryClient:
      """Client for log queries (JSON transport, avoids gRPC overhead)."""
      client = google.cloud.logging.Client.from_service_account_info(...)
      client._use_grpc = False
      return client

For me that is how I know something like protobuf is good. It is a nuisance to manage and distribute the definitions, adds a build step even to languages with no build step normally, is slower than almost every alternative, and artificially restricts you from doing lots of common things. It's so good!

And look at the code quality of the implementation! It's like a team of interns wrote it while drunk. It is a complete spaghetti mess, but has tons of super convoluted micro optimizations that are slower than just doing the most obvious thing, but make the implementation confusing and indirect. It's trash code.

tomtom1337 5 hours ago | parent | next [-]

What are good alternatives when you need a common "single source of truth" schema shared between multiple languages? We use protobuf between c# and Python.

throw1234567891 4 hours ago | parent | next [-]

Json schema?

IshKebab 4 hours ago | parent | prev | next [-]

I quite like the look of Typespec though I haven't used it much.

I always thought Thrift was waaay better than any of the alternatives, but it always had terrible documentation and I think it died mainly because of that.

ltbarcly3 5 hours ago | parent | prev [-]

The goal is not to have a single source of truth schema. That is a means to some other goal, and it's not even a good means.

If you never change the schema then you don't have to worry about it, get things working and never look back.

If you do change your schema from time to time, you need testing between the two systems. If you have good tests again a single source of truth is fully redundant, both systems are talking just fine. If you don't have tests things can and will break all the time even using protobuf.

afavour 5 hours ago | parent | next [-]

> The goal is not to have a single source of truth schema. That is a means to some other goal, and it's not even a good means.

It’s about data transmission. Being able to encode and decode in a type safe manner between different languages (and so, different platforms) is a goal that makes a lot of sense.

> If you do change your schema from time to time, you need testing between the two systems

Or you could just use a defined format that doesn’t require testing. I rarely use protobuf but I can see why people do. The guaranteed backwards compatibility is huge for people who can’t just publish a new web frontend at the drop of a hat.

kccqzy 5 hours ago | parent | prev | next [-]

If you understand how to evolve protobuf schema definitions, then you don’t really need testing. You instinctively know how the parser works when it is parsing data with a different schema from what it expects. And that’s a powerful thing. If your things break even when using protobuf then you don’t grok protobuf.

It’s probably not an exaggeration to say that being able to avoid tests between different systems who have different versions of the schema is a core goal of protobuf. Why? These two different systems are probably owned by different teams, and introducing explicit tests between different versions of them increases coupling between them.

andai 4 hours ago | parent | prev [-]

Both sibling comments say one type of assurance makes the other irrelevant, but I would wager they cover different territory.

onei 5 hours ago | parent | prev | next [-]

Is that a recent-ish improvement? I feel like HTTP/2 would be roughly the same performance for JSON and protobuf, so maybe this is HTTP/2 vs HTTP/3?

ltbarcly3 5 hours ago | parent [-]

I think the overhead is protobuf itself but I can't check.

pjjpo an hour ago | parent | next [-]

Comparing a wrapped C++ gRPC backed stack with an httpx/requests backed one is like comparing apples to elephants.

okanat 5 hours ago | parent | prev [-]

Protobuf simply encode things way more efficient that JSON can define a single object. You're quite frankly spewing bullshit in this whole thread.

throw1234567891 4 hours ago | parent [-]

You don’t get what they say. It’s not about about how efficient it is after encode, it’s about how fast encode is. They are not spewing bs, they’re focusing on a single point. The question is: do you send it over the wite more often than performing encode/decode.

blanched 4 hours ago | parent [-]

That's what the person you replied to is talking about, and they're right. Putting aside the final byte size (where protobuf also wins), protobuf is faster at both encoding and decoding than json. There are numerous benchmarks you can find that show this.

The advantages of json are not related to performance.

someothherguyy 4 minutes ago | parent [-]

doesn't seem universally true, https://github.com/protobufjs/protobuf.js/issues/2114#issue-...

pastel8739 5 hours ago | parent | prev | next [-]

Is this because load is lighter on their JSON endpoints that their gRPC ones?

itsthecourier 5 hours ago | parent | prev [-]

so how do you save data over the cable when it's needed?

throw1234567891 4 hours ago | parent | next [-]

Eventually it’s all bytes. Where do you want to save them?

ltbarcly3 5 hours ago | parent | prev | next [-]

zstd?? Obviously?

anematode 5 hours ago | parent | prev [-]

JSON. /s

echelon 6 hours ago | parent | prev [-]

Buf's offering of protobuf registries and codegen SDKs for microservices seems less necessary in the LLM era.

I'm starting to question many of protobuf's advantages (perhaps not the wire format). Add to that monorepos and other fads of the 2010s given the rise of LLMs.

I used to be a big believer in this stuff, but I'm quickly having my core assumptions change out from under me.

kristjansson 5 hours ago | parent | next [-]

A big differentiator is whether one imagines an LLM in-band with most/all future software. If there is, and we’re deferring until very late parts of a program that world have been load bearing, and we’re able to programmatically ands reliably squint and say “eh i know what you meant” … then yeah formalizations seem superfluous-to-counterproductive.

OTOH if LLMs are to write, but not supplant, much of software, then boundaries, delegation to deterministic layers, good compilers to bonk miscreant models on the head with error message seem essential.

At one point it would have been shocking to assert that the compiler would live in-band with the program too. and yet JS eats the world. It seems shocking today that we could have a universal prior over the world operating in the ms/us nJ/pJ range required. And yet … ?

giancarlostoro 5 hours ago | parent | prev | next [-]

The real argument shouldn't be about protocols becoming obsolete, but programming languages that are "less efficient" could eventually become obsolete in favor of highly scalable and performant languages due to LLMs when the main gap becomes knowing how the tech works at a high level, and not the syntax, why code in one language over another if you don't need to worry about messing up on syntax, only about reviewing logic for sanity and correctness against business rules as well as validating that it is stable code.

cyberax 5 hours ago | parent | prev | next [-]

Anybody who thinks that you can just chuck unstructured data into LLM and YOLO the app is an idiot.

This works up to a point, and then it doesn't. And you're left with tons of inconsistently formatted data.

My company is built on protobufs from ground up :) We use it in the database, for remote calls, on the frontend, etc. The protobuf language is not great, but it's about the right balance between too expressive and too restricting.

And the best thing is that it's compact, compared to OpenAPI.

sudorandom 6 hours ago | parent | prev [-]

I feel exactly that way about REST. The assumption that LLMs make schemas obsolete misses how structured outputs actually work in production. When you have probabilistic models generating code, strict contracts become more critical, not less. It is no coincidence that several major LLM platforms rely on ConnectRPC and Protobuf for their own APIs.

echelon 6 hours ago | parent [-]

You're right, but the adeptness of models to spin up clients and behaviors on the fly is remarkable. They're capturing the semantics of behavior at a deeper level.

If we do strict schemas, I'd like to see less ceremony around them. Tool calls instead of brittle build steps and protocol registries.

Perhaps we need new tools for this going forward.

someothherguyy a minute ago | parent | next [-]

why? are humans going to stop using services?

sudorandom 6 hours ago | parent | prev [-]

Hm... Maybe. In my view, an IDL is part of the input that you absolutely want humans to author or carefully review at least. In my experience, the ceremony around generating code is also performed very well by LLMs. But I do agree, there's definitely some changes that are needed to integrate Protobufs better. Some languages have built-in tooling to make it seamless, but it's definitely not universal.