| ▲ | NitpickLawyer 6 hours ago |
| > But I also think the demand for "fast/cheap/good-enough" models is just about to take off. There's a sort of "revelation" I had in ~early '24 when I used a 7B local model with a library called Guidance (initially out of MS, then the team moved) to create a flow where the model would receive pseudocode for tests, first write the tests, and once I approved then started writing code until the tests passed. This was before "thinking" models, and yet using that library I was able to "guide" the model in the required "prompt / instruct" context such that it was working towards completion, and I saw the first things like we see now in the thinking traces "oh, test x doesn't pass because blah, I need to..." and so on. Anyway, the revelation was "even if the models never improve, I'll have years of fun finding out all the ways I can use these things". And, obviously, the models improved a lot since then. But I think that revelation can still be applied, as a sort of "truism". We have, right now, access to things that 10-20 years ago would be considered magic. We are still finding ways of cobbling together systems with glue, duct tape and prayers and find new things they can do. I think the "good-enough" stage has come not just for API models (cheap, fast, etc) but for local as well. Even if slower, even if clunkier, but they are good enough for a set of ever increasing tasks, and what's more it's incredibly fun to work with them. |
|
| ▲ | swatcoder 4 hours ago | parent | next [-] |
| Yes. The infancy phase of this technology is represented by the pursuit of making wildly grand, wildly expensive, all-purpose models that somehow discern a user's full accurate intent from a lazy, underdeveloped, vague idea that they ambiguously and poorly express in a couple dozen words. The adolescence will arrive as those outsized and ill-considered ambitions collapse and we instead see a cambrian explosion of restrained but efficient model+harness-tuples that have been distilled, finetuned, and rigged to deliver on narrowly scoped but idiosyncratically-shaped tasks with incredible efficiency and erogonomics. |
| |
| ▲ | jimmaswell 3 hours ago | parent | next [-] | | This idea has failed to pan out time and time again - people have an instinct that hand-crafted finely-tuned specialized AI systems must be optimal, but throwing more scale and compute to something more generally smart always wins out. It's especially palpable just looking at the last few years of LLM's: a frontier model with all the world knowledge you can stuff in it and every tool at its disposal has always performed the best at all tasks. Suggesting otherwise has become an extraordinary claim requiring extraordinary evidence. http://www.incompleteideas.net/IncIdeas/BitterLesson.html Recent comment touching on this in relation to LLM's in more depth: https://news.ycombinator.com/item?id=49322695#49323341 | | |
| ▲ | nickysielicki 2 hours ago | parent | next [-] | | The Bitter Lesson is very popular right now. It seems true right now. It’s having its moment right now. That doesn’t actually mean it’s axiomatically true. Commenter below gets it absolutely correct: stockfish, which runs on your 5 year old phone, is dramatically better at chess than Fable. Like, so much better that it’s not even remotely comparable. The theory of the Bitter Lesson, and it’s only a theory, is that LLMs could eventually outperform stockfish. It’s not true today and it remains to be seen whether it will ever be true. For now, specialized models are absolutely better at specialized tasks. | | |
| ▲ | mikepurvis 12 minutes ago | parent | next [-] | | But isn't that really just about giving "front end" models more access to specialized tool libraries, which include models tuned to specific tasks? Like the first model says ah, we're being asked to code something, oh and we've been provided with some example code, let me invoke a tool call to my model the recognizes many languages, that model says that we're looking at ocaml. Okay, I better pass this off to my ocaml model which will decipher the supplied code and make a plan for what we do about the user's intent. The ocaml model recognizes that there are tests in the supplied code, let's have the special testing model have a look at the testing strategy and see how that fits in with what we just implemented, etc etc. And perhaps at the end it all gets a single pass by a god-tier model for overall sanity and congruence, but the actual work, planning, coordination, and even user interaction was done by cheaper and faster agents of much more limited capability. | |
| ▲ | Evidlo 2 hours ago | parent | prev | next [-] | | This seems really backwards. The Bitter Lesson is all about large data-based approaches vs hand-crafted ones, it doesn't say anything about language models not trained specifically for chess. I can't find the comment you're referring to, but the latest versions of stockfish are based on neural networks trained on millions of games, so if anything the Bitter Lesson turned out true here. | | |
| ▲ | nickysielicki 28 minutes ago | parent [-] | | The conclusion of the bitter lesson would be that a large language model trained on chess commentary as well as being trained on millions of chess games would outperform stockfish which is only trained on millions of chess games. There’s no evidence at this point that this is true. |
| |
| ▲ | Animats 2 hours ago | parent | prev | next [-] | | Good point. Dumb AIs are needed for customer service. Most of that industry is still at "press 1 for sales, 2 for billing..." and needs something that will run locally on a 1U server. | |
| ▲ | PEe9bB7D 2 hours ago | parent | prev [-] | | Maybe depends on how you ask it? Directly, or let it write a chess program? I think the latter can yield way better results. |
| |
| ▲ | ZainRiz 2 hours ago | parent | prev | next [-] | | I'd respectfully push back on the framing here. If you look at value as purely the LLM output, then there's a valid argument that the best frontier models will always be better than fine tuned specialists. (I'm not convinced personally, but it's a defensible claim) But that misses two dimensions:
1. The cost of acquiring that output
2. What is actually "good enough" for that specialist domain Not every output needs to be the best to produce value. And as specialist models increase in cost, their cost/value proposition goes down. At some point, there's a threshold where cheaper, fine tuned models are "good enough" at the task and also substantially cheaper than the expert models. That's where fine tuning helps. Personally, I became a believer in fine tuning after fine tuning a 1B Qwen model as a second pass over my local voice transcription app, achieving excellent accuracy at ~zero token cost and waaaay lower latency than if I'd invoked my Claude subscription under the hood. | |
| ▲ | srcreigh 2 hours ago | parent | prev | next [-] | | No. The bitter lesson is about capabilities. GP is talking about efficiency. GP isn’t suggesting that focused narrow model(s) will be more capable than large model, but that many small focused models can have sufficient capability while being more optimal. Also, the bitter lesson is just wrong. The bitter lesson is about hand tuned AI vs computational general methods. However in truth today’s AI uses both. We have general compute heavy models which require narrow expert instructions (eg tools internet docs). LLMs would not be as good without expertly written context, and expert context without LLMs aren’t as good either. | |
| ▲ | joefourier an hour ago | parent | prev | next [-] | | > It's especially palpable just looking at the last few years of LLM's: a frontier model with all the world knowledge you can stuff in it and every tool at its disposal has always performed the best at all tasks. Suggesting otherwise has become an extraordinary claim requiring extraordinary evidence. Absolutely false. At least when it comes to multimodal inputs, even a simple classifier will outperform the largest LLMs who still hallucinate details or don’t describe audio and images accurately. And there’s also the issue of cost/inference speed. Running a trillion parameter model for all tasks will be incredibly costly, require a cloud API, while a tiny CNN can be run locally or at a cost multiple orders of magnitude lower. | |
| ▲ | applfanboysbgon 3 hours ago | parent | prev | next [-] | | This idea has not failed to pan out at all. I work for a startup that is exactly what GP described, and am set for life because of how wildly successful it is. Notably, we are successful, in a genuine sense of the word: we bootstrapped from running tiny models to larger and larger models on our own slowly improving fleet of GPUs, and now have millions in revenue without a single dime of outside investment. Conversely, you cannot call taking on ~1 trillion in debt and purchase commitments to scale "success". OpenAI and Anthropic are underwater financially. To be precise, they're in the Mariana Trench. | | |
| ▲ | wild_egg 3 hours ago | parent | next [-] | | Wait, you actually found a viable counter to The Bitter Lesson? Please say more | | |
| ▲ | klipt 3 hours ago | parent | next [-] | | Perhaps an analogy to Moore's law? Bitter lesson #1: don't waste time optimizing code when a faster processor is around the corner. What countered it: Moore's law stopped working. Bitter lesson #2 similarly relies on scaling laws that might have diminishing returns wrt model runtime vs intelligence. Runtime matters for turnaround on the problem you're solving. | | |
| ▲ | wild_egg 2 hours ago | parent [-] | | Moore's Law has nothing to do with processors getting faster. Dennard scaling stopped working but Moore just slowed somewhat, not stopped. |
| |
| ▲ | applfanboysbgon 3 hours ago | parent | prev | next [-] | | This is a misunderstanding of either the bitter lesson or what was being claimed, on multiple accounts. Firstly, the bitter lesson is merely about human expertise-tuned algorithms vs. throwing raw compute at a domain. But, notably, it is still domain-specific. No matter how much compute you throw at training an LLM, it is never going to beat a Chess engine at Chess. If you give a Chess engine 1,000,000 compute units and a general-purpose LLM 1,000,000 compute units, the Chess engine is obviously superior at Chess; ergo, there is value in throwing compute units into training models for specific tasks. This is true for within several orders of magnitude of compute, in fact. It's also true that if you give the Chess engine 1000 compute units it'll still beat the all-purpose model with 1,000,000 units, so actually there's a lot of value in training for specific tasks. Secondly, the bitter lesson is predicated on compute being cheap. There was a period where a hand-tuned algorithm informed by human expertise would outperform a raw alpha-beta search at Chess. Then compute got cheaper, and DeepBlue ascended to the top. Compute is now expensive again relative to the tasks being performed. We are absolutely still in a period where human expertise in training LLMs will outperform a naive approach with more raw compute. | | |
| ▲ | CamperBob2 2 hours ago | parent [-] | | I don't know much about chess engines; do they still use hand-tuned algorithms, or are they more like AlphaZero, where they learn through self-play to beat any/all possible human contenders? I don't believe DeepBlue was automated to that extent, but it may have been. In the latter case, the chess example would tend to support the Bitter Lesson, rather than refute it. I would also be VERY slow to claim that general-purpose models will never be competitive at chess. It wasn't so long ago that transformers couldn't add two-digit numbers reliably without resorting to tool use. They are now as good at "mental arithmetic" as any human savant. It wouldn't surprise me at all to see someone come up with a model that just happens to be really, really good at leveraging the portions of its general training data having to do with chess. In fact you could argue that AGI demands such a model, if we are to assume that LLMs are a guidepost in that direction. | | |
| ▲ | dmoy an hour ago | parent [-] | | I don't know anything about the last 8 years of chess engines, but yea maybe 8-10 years ago AlphaZero shit all over e.g. stockfish. |
|
| |
| ▲ | HDThoreaun 3 hours ago | parent | prev [-] | | The issue is that GP is misusing the bitter lesson. Yes, search + learn tends to be more effective than human rules based strategies, but that's not what's being considered here. The original claim is effectively that AGI isn't needed for most tasks and more value can be created by using search + learn to solve specific problems instead of applying general models to every problem. Then GP commented a non sequitur |
| |
| ▲ | z3t4 3 hours ago | parent | prev [-] | | Do you have a website? |
| |
| ▲ | CamperBob2 2 hours ago | parent | prev [-] | | Suggesting otherwise has become an extraordinary claim requiring extraordinary evidence. VibeThinker 3B constitutes extraordinary evidence, IMO. The first such evidence I've seen myself. Very small model, very low literacy, almost no world knowledge, but it is as good at math and logical reasoning as models a hundred times larger. The Bitter Lesson is a valid and trenchant observation about how about we got here, but I think it's a mistake to assume it tells us very much about where we're going. Too much has changed recently and is still doing so. | | |
| ▲ | spockz 2 hours ago | parent | next [-] | | So theoretically, if you give that model the means to find information, ascertain the quality of said information, it could still reason its way to an proper answer? Is this whole thing than maybe a read vs write optimisation again? Spent more time and effort training more knowledge into the model upfront and get it out in a single question instead of training a small model and needing more steps to answer the same question? | |
| ▲ | algo_trader 2 hours ago | parent | prev [-] | | > VibeThinker 3B constitutes extraordinary evidence.. math and logical reasoning Any similar model aimed at coding? A >10B model for mass spawning/swarming and reporting back to a larger model | | |
|
| |
| ▲ | HoldOnAMinute 3 hours ago | parent | prev [-] | | Someone will eventually figure out how to package it all into a single, cheap chip | | |
| ▲ | bmitc 3 hours ago | parent [-] | | That you can then write text to program and make applications with. |
|
|
|
| ▲ | jermaustin1 5 hours ago | parent | prev | next [-] |
| To me, most local models work just fine for anything you can be patient for. If I want something quicker, I will go to a SOTA model via API, but with multiple 3090s, I have never really needed a hosted model for a lot of my experiments. For code, they are great, but for creativity for NPC controllers, they leave something to be desired, but work well enough for testing, so I don't burn tokens until I'm actually playing my games. But nothing one-shots a prototype better than Fable 5. I can have a prototype built in 30 minutes, hooked up to my local LLMs and Claude Code is very good at testing the interactions and even tuning the prompts of the NPCs for better experiences. |
| |
| ▲ | __float 5 hours ago | parent | next [-] | | "with multiple 3090s" is quite a bit of burying the lede for "most local models work just fine", don't you think? | | |
| ▲ | jermaustin1 5 hours ago | parent | next [-] | | Having multiple 6 year old cards doesn't seem like it's that big of burden for local LLMs. I get that a lot of people don't have them. And a single one can be VERY performant. And the smaller models like a 7B can run on much smaller hardware like a mid-range [3|4|5]060. My entire AI Dev Box cost $4500 in parts. 128GB RAM, i7-10700, 1TB and 2TB SSD, and 2x 3090s. Today's prices and inflation have definitely made that price tag seem a lot better than it was, but it was an investment in all things GPU that were happening in 2020 (crypto, blender, image gen), then LLMs exploded. | | |
| ▲ | thayne 5 hours ago | parent | next [-] | | A single, used 3090 costs more than I have ever spent on a computer. | | |
| ▲ | 9cb14c1ec0 4 hours ago | parent | next [-] | | Yes, the tunnel vision around local models on this site is crazy. The percentage of people in the world who can afford the hardware is extremely low. | | |
| ▲ | layer8 4 hours ago | parent | next [-] | | It seems roughly similar to the pricing level of personal computers in the early eighties (i.e. IBM PC and Apple Macintosh). I’d expect prices to come down significantly over the next few years. Not so much in the next year or two, but after that. | | |
| ▲ | oblio 3 hours ago | parent [-] | | Just like for warships, the complexity and cost of building cutting edge hardware has grown exponentially up to a point where a significant chunk of the world's computing is dependent on 2 companies: ASML, TSMC. We shouldn't extrapolate linearly from examples from the 80s. | | |
| ▲ | layer8 2 hours ago | parent [-] | | No, but I wouldn’t expect it to stagnate like with Intel in the 2010s either. Maybe the biggest caveat is that most people will be fine with using cloud providers, so the market for non-server hardware won’t be subject to as much competition. |
|
| |
| ▲ | regularfry 2 hours ago | parent | prev | next [-] | | There's a non-small contingent who lucked into the periodic games machine upgrade at the right time to snag a {3,4,5}090 rig just before everything exploded. It's a small contingent now but it was less so then. And now those people can add a second card for roughly what that whole system would have cost new originally. | | | |
| ▲ | jermaustin1 4 hours ago | parent | prev | next [-] | | I don't think there is tunnel vision. I'm just saying that I have a couple 3090s I invested in a handful of years ago, and they are still going strong today as multiple GPU-needing technologies emerged. I'm not saying everyone has to run local LLMs, because the APIs are in a race to the bottom, and my $10 of OpenRouter credits I bought months ago is down to $8.94 because most models give you MILLIONS of tokens for a US Quarter. | |
| ▲ | BoxOfRain 3 hours ago | parent | prev | next [-] | | It's a decreasing pool as well I'd say, the dev machine I built last summer would make less financial sense to me now for example. | |
| ▲ | zahlman 3 hours ago | parent | prev [-] | | I mean, I'm not rushing out to buy that kind of hardware myself, but it is a matter of perspective. People commonly spend an order of magnitude more on a car, and that's just the sticker price. |
| |
| ▲ | wafflemaker 4 hours ago | parent | prev [-] | | My single 3080 runs so hot I don't need to warm my room in winter, and have to play games in my underwear in summer. |
| |
| ▲ | zamadatix 4 hours ago | parent | prev | next [-] | | I got a great deal on ~72 TB of NVMe right before storage prices shot up, doesn't make it any less ridiculous that I have it or any more relevant to people talking about building a NAS now. 99% of people, even in tech, do not have the stupid amounts of hardware people like us hobby on. | | |
| ▲ | oceanplexian 3 hours ago | parent | next [-] | | Most people in the US have a car, and the average new car is $40,000. Hell where I live a middle class consumer will spend double that on a Boat or an RV and think nothing of it. These aren’t elite tech workers. It’s not unfathomable that if a personal, generally intelligent local AI provides enough utility and doesn’t require you to tweak CLI flags millions of Americans would want one. | | |
| ▲ | spockz 2 hours ago | parent | next [-] | | Spending that kind of moment on a product that gives you personal happiness for years up to decades and then will still have residual worth, which people save up for ages for, is an entirely different proposition than buying a product that may make you faster professionally, but which in the short time can also be achieved by a few dollars worth of subscriptions to a hosted model for even greater effect. | |
| ▲ | ninglor an hour ago | parent | prev [-] | | Most people in the US don't drive a new car, and used cars can be had for far less than $40k. An $80k purchase would be just shy of the median annual household income -- anyone who thinks nothing of that has financial resources far above typical. You are in a bubble. | | |
| ▲ | vel0city 7 minutes ago | parent [-] | | An $80k purchase is far more affordable when you're looking at an 84 month loan. You trade in your current $20k truck with $30k in debt on it for your $80,000 car, get a couple grand in incentives and a $10k down payment, and boom you're only looking at a bit under $1,200/mo in payments. The median household is bringing home ~$84k before taxes, hypothetical person lives in a no income tax state, they take home ~$5k/mo. Easy peasy, its not like you were planning on taking any vacations anyway since you're always working. What matters is you've got the Duramax HD King Ranch TRD Big-Boy machine. Doesn't matter the cost. You can tow anything, drive anywhere, do anything, and do it all in comfort. Other than parking in a normal parking spot comfortably. Or even park it in your own garage at home. I've seen this exact scenario many times personally. |
|
| |
| ▲ | sroussey 4 hours ago | parent | prev [-] | | where? i would love that. | | |
| ▲ | zamadatix 31 minutes ago | parent [-] | | "Where'd I buy it" or "where is it now" ;)? It was a 96 core gen 4 epyc+supermicro board build with consumer NVMe drives on 1x16->4x4 "dumb" bifurcation cards. I had to get a few MCIO-> PCIe adapters as well to get the full lane coverage. Mounted in a standard EATX compatible consumer case with a consumer PSU and a lot of Noctua fans - surprisingly cool and quiet for what it is. Motherboard+CPU I got from Ebay. Rest from the best MicroCenter/Amazon/Walmart deal of that day. Bought juuuuust before the AI pricing apocalypse, largely by pure chance. |
|
| |
| ▲ | xnx 4 hours ago | parent | prev | next [-] | | > 2x 3090s You could sell those and have enough money to pay for hosted inference for years. | | |
| ▲ | jermaustin1 3 hours ago | parent | next [-] | | They cost more to run than hosted anyway. But that isn't the point of having them. They are a playground, a backup when the internet is down, or claude is down. They can render Blender scenes pretty well. They play any game I want. You can do each of those at various hosts and own nothing. Or own a couple "over priced" cards and do it all at home on battery power for a few hours while the power is out. | | |
| ▲ | robotresearcher 3 hours ago | parent [-] | | For me it’s more that you can show them your financial and medical data without BigCo looking over your shoulder. |
| |
| ▲ | Gecko4072 3 hours ago | parent | prev | next [-] | | But after all those years you’d still have 2 3090s, which are now about 6 years old and still holding value. | |
| ▲ | irishcoffee 3 hours ago | parent | prev [-] | | I keep seeing this comment. This is _hacker news_ where, back in the day, people just hacked on things, because it was a hobby. They weren't "moneymaxxing" or desperately trying to be as insanely efficient as possible. They hacked on stuff with a can of surge at 3am because it was fun. Your comment is like a meta comment of "LLMs are generating everything, after a while the ouroboros will eat itself. (Which I agree with)" If people aren't hacking on this shit just because, you have completely conceded control of software to a handful of sociopaths, and open source software is dead. |
| |
| ▲ | vel0city 16 minutes ago | parent | prev [-] | | [dead] |
| |
| ▲ | bitexploder an hour ago | parent | prev [-] | | Not really. 2 years ago that was a pretty normal amount of GPU hardware for a hacker or gamer. It's all relative. They are not accessible to most people yet, but for someone that cares and is a technologist? Likely accessible. |
| |
| ▲ | sroussey 4 hours ago | parent | prev [-] | | I have trouble getting simple extraction to work sometimes. I have a block of text describing people and their roles at a company and their ages, and i asked for structured results of an array of these things with the text span that it appears in and all i can say is: nope. |
|
|
| ▲ | keeda 2 hours ago | parent | prev | next [-] |
| Yep, I've been having excellent experiences with the models even from the 2023 era. They required a lot of "holding it right" (mostly: being very precise in what went into the context) but their raw coding capabilities were astonishingly good even then. However, back then I was getting the AI to write individual functions or classes or a test suite. I was decomposing the larger task into smaller tasks, delegating some of them to the AI, reviewing the results and composing the codebase from those. I was also essentially the harness. Today the models can write and test and deploy an entire project. In terms of the code quality, I actually don't think today's frontier models would have written it much better than the 2023 models did. So in terms of raw coding capabilities i.e. converting a high-level specification into working code, I think we hit the peak way back in 2024 itself. What has changed is the AI has learned how to do the task I was doing (besides being the "harness"!), which was the mid-to-higher level "engineering" aspects like decomposing a task, specifying it to a reasonable level, reviewing the outputs, and course correcting as needed. I'm not sure if that is something the AI labs explicitly focused on during training (which may be why Meta is having its highly paid engineers do annotation work), or an emergent property of "better reasoning" (which I believe Dario implied in a podcast), or some mix of both. But the fact remains that even the weaker models are more capable than we realize, and many being open weights, are here to stay. |
|
| ▲ | ksec 4 hours ago | parent | prev | next [-] |
| While they are improving rapidly, or as you say even if they don't. The next stage is for hardware companies ( cough Apple cough ) to ship these Local Model ready hardware in their products. It will be interesting to track the improvements of these 7B model over time. There will be a turning point in the next few years where it attract enough consumer attention to create yet another Smartphone and PC super cycle. |
|
| ▲ | gozzoo 35 minutes ago | parent | prev | next [-] |
| > We have, right now, access to things that 10-20 years ago would be considered magic These things would be considered magic even 4 years ago! |
|
| ▲ | nowittyusername 2 hours ago | parent | prev | next [-] |
| There's A LOT low hanging fruit still out there for sure. And with antigenic systems being able to do the boring repetitive work of looking for that low hanging fruit I think we will see interesting things indeed. Also I think heuristics is where its at for such things. Once you describe some good heutistical structures for the research models to always follow related to "creativity" and such things, thats where we will see biggest difference. The agentic systems know the scientific method well and can follow it they just need the ability to be "creative" so their sampling becomes less rigid. |
|
| ▲ | eqmvii 4 hours ago | parent | prev | next [-] |
| I see it in a slightly opposite way: even the good models are relatively cheap, and so I worry what we might miss by spending too much time playing with the Sonnets of the world when the Opuses are still objectively a bargain for the power they bring. |
| |
| ▲ | zahlman 3 hours ago | parent [-] | | > when the Opuses are still objectively a bargain for the power they bring. The cost isn't just what you're billed. There are security, privacy etc. concerns. | | |
| ▲ | Foobar8568 3 hours ago | parent [-] | | I know companies that are using github, even using public repo, and request their teams to not use SOTA models, but are ok with local models.
Just stupid policy. |
|
|
|
| ▲ | 5 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | riazrizvi 4 hours ago | parent | prev | next [-] |
| I think there's something subtle about language and ambiguity that means they aren't designed to become superintelligent autonomous machines. They're value is as information repositories that actual intelligent autonomous machines (us) mine and string together. |
| |
| ▲ | dgellow 4 hours ago | parent [-] | | Yes LLMs are a beautiful way to compact knowledge. It would be such a cool technology to develop and worked with if it wasn’t linked to such a toxic industry | | |
| ▲ | riazrizvi 3 hours ago | parent [-] | | I think you're just observing ppl in one of these rare instances where enough of them come together because they are motivated. 'Toxic' is the clamoring sound of a crowded room where what gets through to your ears are just the most annoying snippets of incomplete conversations. I dare you to hang out with any actual people here, understand their viewpoint and listen to what they actually have to say in person, within the context of watching them do it. | | |
| ▲ | dgellow 40 minutes ago | parent [-] | | I know those people. Lots of them are fantastic humans. That doesn’t change the fact the AI industry is extremely toxic |
|
|
|
|
| ▲ | QuercusMax 28 minutes ago | parent | prev | next [-] |
| Just being able to instantly generate a complicated query expression to pull specific bits out of a JSON blob sold me. It's awesome that I can ask Claude to build a whole feature and it will often one-shot it for me, but generating utility bash / python scripts or little throwaway utility webapps is what really excites me. |
|
| ▲ | LoveMistral 5 hours ago | parent | prev | next [-] |
| Same. Mistral 7b has been more than I ever needed for text for years now. Unless you must 1-shot with no harness it’s the same amount of power, maybe more because the big “good” models make too many assumptions and tend to become rigid. Mistral 7b can do anything, and it’s basically instant even on an M3 |
| |
| ▲ | frigidwalnut 5 hours ago | parent | next [-] | | Sounds interesting. Can you give more details on your workflow and what tasks you use it for? | | |
| ▲ | LoveMistral 4 hours ago | parent [-] | | Code, creative writing, email summaries, automated email replies, and I prefill my invoice notes and daily updates for work. Actually built a full invoicing product for that, using it too. I use Mistral 7b and LlamaIndexTS on Node, I run it on a MacBook M3 and on a Linux server with only 8GB VRAM (old gaming PC). Basically flawless, runs very fast and I don’t even know what paying for “tokens” is :) | | |
| |
| ▲ | Almondsetat 5 hours ago | parent | prev | next [-] | | What kind of work are you doing? For example, if I have some code in the hot path and I want to do all the usual tricks to help the compiler vectorize it, such a small model is not able to do much. | | |
| ▲ | LoveMistral 4 hours ago | parent [-] | | RAG is your friend (or any vector db). No model can vectorize an entire codebase in context. Even a big mainstream product (like Gemini) cannot handle more than ~1k lines without missing details and making mistakes. And about every 1k lines, it seems to forget the previous 1k, doesn’t it? So you can never hold more than a file or 2 (or 3) in context at a time without losing details. What you find is that the big models like Gemini are doing vector storage and retrieval too, and breaking prompts down into chunks for various models to handle to assemble a thorough response. If you want that kind of control in your outputs, and be able to hold a lot in your inputs, I don’t see any other way regardless of which model you use. | | |
| ▲ | usef- 40 minutes ago | parent [-] | | Out of interest, have you tried the newer models? You are not describing my experience recently. |
|
| |
| ▲ | casper14 4 hours ago | parent | prev [-] | | What are some limitations you have found with using a smaller model like that? |
|
|
| ▲ | viscousviolin 4 hours ago | parent | prev | next [-] |
| If someone has an old GPU laying around, say a GTX 1080 with 8 GB of memory, would that be enough to get a (small?) local model running? |
| |
| ▲ | bityard an hour ago | parent [-] | | A small model, yes! But not necessarily a good model. With the additional caveat that I don't know whether that specific card is supported by modern drivers. You'd be looking at one in the 6B or 7B parameters range at FP8. Or smaller. It's been quite some time since a recognizable company in the AI space released a model that small. You can try larger model that has been quantized down to that size, but they don't always fare well with that. Modern text-to-speech and speech-to-text models also fit well into modest amounts of VRAM. |
|
|
| ▲ | Der_Einzige 3 hours ago | parent | prev | next [-] |
| BTW structured/constrained generation has so many places to trivially enable jailbreaking/alignment/safety problems that closed source models heavily limit the full expresivity of grammars and capabilities, particular of on-the-fly dynamic grammar construction/reconstruction. |
|
| ▲ | dominotw 3 hours ago | parent | prev | next [-] |
| ppl keep talking about the supposed unexplored and untapped "model overhang" but very few things in the world are where you can write elaborate test criteria to before using ai. A sales person sending a prospect email doesnt have a way to write a test harness for it. Yet these tasks dominate what humans do compared to writing a crud app . otherwise anthropic wouldnt have trillions dollar valuation |
|
| ▲ | cyanydeez 4 hours ago | parent | prev [-] |
| I've amassed access to 4 different GPU rigs with 128GB to 72GB; I didn't this before I event touched an agentic engineering harness. It was sometime in February/March when I set them to first tackle small problems, and now with deer-flow, they're scaffolding full project/scope implementation and I'm finishing off the fine details around the problematic edges. |