| ▲ | Ask HN: What default model do you use and why? | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| 59 points by stikit a day ago | 99 comments | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I use claude for most of what I do, and Fable is largely overkill for me and frequently burns through my Max plan's session credits in minutes (!) when just doing an initial mobile app planning with 4 agents. After I waited out the timeout period 6 hours later, and I picked up again, the cache had timed out so it burned through 2% of the session in less than a minute. Opus 4.8 is now my goto and I will be avoiding 5 until I see a reason to switch back. 4.8 is 'good enough' for what I need and has been a great value. It mostly gets things right. Most of what I do is web and mobile , largely cloud backend. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | o_m a day ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I used to main Claude, but I can't stand how it writes. I feel like I'm wasting too much time trying to decipher the output. Adding writing rules does not seem to work. Now I use GPT 5.6 Terra high fast mode, with Luna for everything else. I might consider using Sol for planning. I can't stand using Sol or smarter models for coding, because they will eventually try to rewrite everything in the codebase. I also don't want to use the Claude Code and Codex agent harnesses. The good thing with Codex subscription is that it can be used in other harnesses, unlike Claude. As far as I know, only Anthropic has this restriction. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | montroser a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
It pains me to read these answers so far. Listen, for 99%+ of your web and mobile tasks, deepseek-v4.1-flash is all you need. It is blazing fast, super cheap, and quite proficient. It acts responsibly, has top-notch vision for evaluating its own UI work, and is far less smug and flowery than any of the Anthropic models. For what it's worth, here's a take on its speed vs cost vs intelligence: https://artificialanalysis.ai/models/deepseek-v4-1-flash I can go all day and night with this thing with multiple sessions going, and I spend like $2 per day retail. With opencode-go, that fits within the $10/mo subscription, so that's what it ends up costing in real life. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | hgoel a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
For personal needs I use a local Qwen3.8-Next-Flash setup on a GB10 cluster. For work, Github Copilot with either GPT 5 mini or toss up between Opus/Sol depending on the complexity of the task. Used to pay for a Claude 20x plan and did everything in Opus, but I hate how it talks now and recent events (OAI scooping, Anthropic's spying, third party Chinese model hosts stealing and selling credentials) have really pushed me towards local AI for personal needs. Am not allowed to use Chinese models for work even if self-hosted so not much choice there. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | vallerie a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I've found good success with the new Gemini models on Antigravity. Granted I use my models either: - like a fancy auto complete (here are some stub methods, they should do X, fill them in) - using fairly detailed plans and test harnesses, so blowing up the world is hard The 3.X Flash family have been fairly capable models, and the selling point for me is just raw speed. Gemini is noticeably faster than the competition, about 3-4x, and I just get work done faster with it. That said I'm keeping an eye on Open Weights. DS4 Flash was good until price hikes, and finding a provider that serves at high speed and without quantisation at the prior price is tricky. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | asd88 a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I have “unlimited” (but metered) tokens at work. I was primarily a Claude (Opus and Fable) user until Astra came out. Astra’s outputs are much more concise while being similarly accurate, they seem more information dense. It’s also extremely fast and doesn’t have to “go check instead of answering from memory” if the answer is already in context, or dig super deep into a project before answering. It’s honestly refreshing to work with Astra after being basically burned out from reading Claude’s responses. It’s not perfect. It’s just as “mid” for professional SWE work as Opus/Fable, regularly making incorrect assumptions/generalizations, requiring steering in large codebases, and being incapable of making reasonable long-term software design decisions on its own in complex projects. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | ulrikrasmussen a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Almost exclusively Fable at xhigh. I sometimes run out of tokens, but most weeks it's not a problem. The limiting factor is by far how fast I can push understanding of what was made into my brain, and I find that I waste time using Opus. It makes annoying mistakes, and it makes assumptions that are often wrong instead of actually verifying. And it can't write so I can understand it. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | beej71 a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Given the tremendous variety of answers here, one has to wonder how much it matters. Reminds me of when you ask a bunch of motorcyclists what oil to use | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | hmokiguess a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I think for me stuff peaked around Opus 4.7, I was leaning heavily on the model with paired supervision from reviewing the output manually every step of the way. Ever since that things got a little more complicated and in an unsustainable pace for me, I am trying to remove myself from the equation and build verifiable and reliable tests with quick feedback loops that let frontier models run autonomously but in all honesty not seeing it scale well, I need to take a step back and reassess if the trade off was worthwhile. Frontier models are being incredible at making me feel like they passed my tests only to eventually reveal some tech debt that forces me to take large pivots. It seems the speed is sexy but the results are questionable, models may have hit a limit in my workflow and I think harness engineering is more important than anything. Would love to hear feedback on this take and if others have experienced similar things and what they did to overcome this. (Context would be solo founder bootstrapping greenfield work with full autonomy and sometimes more room for rapid iteration) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | wdm0006 13 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I use a lot of different harnesses, so in CC i use sonnet as main w/ opus advisor, in codex i use astra all the time, in obvious I usually use the GLM models. In general I've found that having a fast main model with a big smart advisor is a nice UX if I'm going to be on keyboard interacting with it, for a long-running agent that I'm not attending to I care much less about the main model speed and more about cost/quality. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Jordan-117 20 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I subscribed to ChatGPT for a few years, but over the holidays noticed that it was making frequent errors and hallucinations when dealing with more obscure Linux and Firefox issues. Trying Gemini as a backup produced much better answers, to the point I switched my subscription. I do everything through the app/website rather than the API; I understand this is pricier but I like having convenient access, a searchable history, memory, etc. I try the ChatGPT free tier as a backup second opinion sometimes, but find their newest model to be painfully rambling. I want to like Claude and might consider subscribing to it instead, but am put off by reports of its tight usage limits. With Gemini, though, I never hit a limit unless I conduct a few Deep Research queries at the highest level. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | time0ut a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
If I had to pick just one, these days it’s Composer. It is just a cheap, fast, surgical workhorse. My workflow uses other models as well for various phases: Opus for planning, Sonnet for review, etc. I prefer Cursor at this point just because of Composer. Claude Code is passable but the lack of a good, fast, cheap workhorse sucks. Sonnet and Haiku aren’t it. I also do like Codex and Sol, Terra, and Luna. They are decent but I don’t find they stand out enough to use over the others. Additionally, I have not tried Astra and found Fable to really not worth the cost for the tasks I do. Finally, the latest Grok is actually a beast of a model, but expensive enough to not be a stand out. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | m0rde a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Enterprise pay as you go across a few vendors. But still cost conscious. Defaulting to Sol Medium/Light for most planning and implementation I think is tricky or want more care in. Luna Extra High for everything else (implementation, tedious take over my browser and do stuff). I read all of its output tokens and lots of thinking tokens to understand the general flow of things, but only minimally look at code these days. I can't grok what Claude models speak and it's gotten worse. OAI models speak my kind of tech language I guess. Light human review, some automated review. Most work is for internal use. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | joelboersma a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
For work, Opus 4.8 because: - I'm on the base Claude Team plan, so no Fable access - Opus 5's outputs are wayyyy to verbose, to the point where I dread using it - Opus 4.8 is intelligent enough for my needs and doesn't make me dread reading Outside of work, I'm using Codex's free tier with 5.6 Luna-High. My personal/volunteer projects are simple enough that I don't need much more. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | codazoda a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I use Sonnit for my personal work and Opus for my professional work. For my personal stuff, I'm on a small $20 plan, so I need to use tokens conservatively. I was very rarely exceeding limits until I built a Dark Software Factory. It's not as efficient at token use. So, I use Sonnit over Opus here. At work I have a $100 plan that I rarely exceed so I use Opus. I have access to Fable too, and I did use it a lot while it was new, but I don't find it improves most of my work by too much. I do mostly bug fixing across several hundred repositories with hundreds of thousands of lines of code, mostly written by humans over the past 20-years. These projects interact with each other so I run claude from the root of my sandbox (I was nervous to try this but I'm not looking back now). I also use Sol as a secondary for my personal work. I pay for it because I like to talk to ChatGPT on the web. Since I already have the subscription, I let Sol write plans for me. It does a better job at certain tasks and it saves me some Claude tokens. Maybe I should consider Terra for the task, but I don't run up against my usage limits for the little bit I use it. I'm trying to use Gemma 4 12b for some workloads but I haven't mastered the model yet. It's still very experimental for me. I can get work from it but it takes a lot of hand-holding. For local models, however, it's all I have the RAM for. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | seanmcdirmid a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Gemini 3.8 flash via Antigravity desktop has been my daily driver for the last few weeks. I also use DeepSeek V4 flash and a Qwen 3.6 MoE for local stuff. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jakevoytko a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
My workhorse is Opus orchestrating Sonnet or Sol orchestrating Luna, depending on which ecosystem I'm using this week. I'm a backend engineer, so for tasks that either interact with the machine learning or client codebases, I have Fable (Medium) review across the product spec, tech spec, and all the codebases after the implementation is done (haven't tried this workflow with Astra yet, don't know if it's a good reviewer). As a side project, I'm doing an experimental task now (having Astra implement my own personal Google Docs clone, OT and everything, for my blog), and it feels more capable than Fable but needs to be watched closer than Fable. It is obsessed with verification and evidence to the point that I've needed to make some absolute rules to stop it from repeatedly e.g. making unit tests with 10,000 test cases with 5+ hours of runtime For well-defined coding tasks, I'd default to either Luna or Sonnet, or hell, just writing it by hand. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | george_stephen a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I use all of em, ChatGPT 5.6luna ion web, for quick tasks(without signup), (ii personally feel ChatGPT isn't straightforward especially when quizing its opinion on a topic) , Gemini for academic work to encyclopedia and know-how discussions, recommending solutions to technical tasks by looking up web for me(it's my primary driver, but lags behind in coding abilities,its large context is a plus. I recently started returned to claude and it isn't bad, since I'm not a paid user... So i don't have a personal assessment on newer models capabilities. Grok is keeping up and really good at explicit stuff, i think that's a drawdown for me. I think most nsfw was made with grok As for meta ai, i feek it's a total joke, i don't have access to muse cause meta has decided not to support older OS versions (I'm on Android). This comment box makes me feel almost as if I'm coding, because of the font. Pretty noice | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | gtadesktop1 a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I used Claude too, but I was banned so I searched for new like you yet and found Google AI Studio. It has a Playground-Mode for the newest models you usually have to pay but in the Playground-Mode you don't pay one cent. You can use Gemini 3.8 Flash and you can give him system instructions. For example you can say "You're a AI Agent that helps me with my mobile app" and he helps you with your mobile app. It has a context window up to 1 million tokens and it's fully free. If the limit is reached you can only click on the "+" icon on the top right and say him what you did attach all project files and tell him what you need. He reads the files you attached and helps you like the old chat. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | alstonite a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Astra low for pretty much everything gets me through the work week with a bit to spare on the 5x max plan. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | expedited123 a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Qwen 3.8 flash-next or Gemma 26B because I care about responsible and sustainable usage of LLM's. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | conradludgate a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
For personal use, Luna-xhigh on the Plus subscription. I only ever hit my limits if I decide to use Sol or Astra for a bit more quality. For work, given we pay API pricing anyway, I've been happy experimenting with K3-medium as my default, and now I'm planning on trying GLM 5.3 as well. Sol-medium is my fallback for my work purposes but I occasionally use Opus 4.8 as an additional reviewer. I like the open weight models because much of my time at work is spent on security hardening (specifically, hardening my own service), and Sol/Opus keep snitching on me and blocking my prompts. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | gpugreg a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I currently use DeepSeek-V4.1-Flash and the V4-Pro and -Flash versions before that, because they are extremely cheap. I have spent less than $25 for over a billion token so far (1B cached, 15M out, 19M in). I even prefer the DeepSeek models over the older OpenAI offerings. When I tell a GPT-5.x model to do some difficult task, they often give up saying it can't be done, or cheat by modifying the tests, while the DeepSeek models are more persistent and less prone to cheating. GLM-5.3, GLM-5.3-Flash and Kimi K3 are also fine, but slower, more expensive, and less good for what I use them for, which is mostly Python and CUDA programming with some JS and HTML inbetween. But almost all of my tasks are verifiable tasks, which means that the LLM can check whether it is done or whether it needs to keep trying. If you are mostly working on problems where the quality metric is based on vibes, YMMV. I haven't tried Astra or Fable yet, because I am not made of money and am happy with my current setup. Also, the Opus models' writing is absolutely insufferable. My blood pressure rises every time I see a Claude-generated slop README. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | mariocesar a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I have an 'ask' alias in my shell that just uses Haiku. I use it daily for pretty much everything For more long work, I now use Fable to create a PLAN.md. I tell it to make a plan that will be executed by other models, and most of the time it ends up choosing Opus or Sonnet. I didn't start doing this recently. Before that, I would just use the top model for everything. Splitting the work across different models depending on the task has helped a lot. They run faster, and I usually get much better results | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | oduis a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I like AI more on the short leash, giving it specific agents tasks, one at a time (centaur mode). I found GPT 5.6 LUNA to be astonishingly capable. Plus, it so fast, that it does not block my flow of throughts, like the more capable but slower models often do. And it is so cheap that I do not use Ollama local models anymore. Luna is far more capable, and so cheap that the energy prices here in Germany eat the gains of local hosting ;-) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | ricardobeat a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Opus 5 at low effort is quite good at most tasks, and will go a long way even on the base plan. At higher thinking levels it is unbearably cautious and verbose - it drives me mad. I don’t have a use for Fable at current prices. My current go-tos are DS Flash, Minimax M3 for subagents. Adding Muse Spark 1.3 to that mix but not convinced yet. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | the__alchemist a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Sol High-Xhigh, and Opus 5. Granted, yesterday I threw a few tasks to Astra which the former 2 botches; it produced clean, correct solutions quickly, so pending further eval, this may take over. IMO unless it's a mechanical tasks, it's worth it to use carefully -crafted queries on the more expensive models, than iterate through messier solutions on the cheaper ones. edit: maybe not. Astra has imitated Fable, and appears unusable for biology due to safeguards. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | karmakaze a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Personally I'm using Qwen3.8-27B (MXFP4 quant W4A8) locally hosted on a pair of AMD GPUs (with DeepSeek Harness). It starts at 250 tokens/sec down to 120 past 128k context. At work mostly Opus 4.8 (sometimes a GPT or Gemini 3.1 Pro). I find Opus 5 chatty/slower and Fable can venture into over-engineering itself into unnecessary complications. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | verisimilidude a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
It really depends. For high-level strategic conversations, I use Fable. For planning, I use Opus or Sol. Sol is generally preferred; it’s faster, cheaper, and less verbose. But I still find Opus more capable on the most nuanced or complex tasks. For planned implementation, I use Sonnet. For one-shot unplanned implementation, I’ll use whatever model seems best for the task. FWIW, I’m increasingly turning to Grok here. I use Luna all over my workflow for reporting. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | traverseda a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Claude, not necessarily because it's better. I find deepseek flash to be of similar quality. But because it's so heavily subsidized. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | sourcecodeplz a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Muse Spark 1.3 Contribs unbeatable price/intel ratio per M tokens: $0.10 (input) $0.20 (output) $0.002 (cached-input) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | shelled a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ZAI/GLM because a friend has generously shared an API key. They have other subscriptions and also a lot more from their employer. A day ago I activated Google AI Pro free via Google's tie-up with a local company (I do pay for this company's product though and it's anything but costly). Now I will use this too. No other reasons to pick these, or not picking anything else. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | prism56 a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I don't deploy anything, I don't make money. V4 flash latest has been great for me. It's so cheap it's almost free. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | seabrookmx a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Sonnet. It makes my token budget go further and I prefer to have the LLM churn away on rote work while I do the thinking. So while Opus and Fable are undoubtedly smarter, that doesn't materially affect _my_ workflow. I'm not as up to date on the other vendors' models, but when I last used Gemini my feelings were similar between Pro and Flash. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jinnko a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I was using various open weights models until glm-5.3-flash came out recently. It's incredibly capable and cheap, even if it's very verbose and not the fastest. I've assigned it to all my agents across my harness and it's getting the job done. Still needs a good steer every now and then, but a great work horse. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | rsyring a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Context will be important for answers to be meaningfully comparable: - LLM expense budget - What type of dev: work, personal, real time spaceship thrust vectoring, html contact forms for family, etc. - human in the loop with short as possible turns, software factories that can run for days, or something in the middle Just off the top of my head. I'm sure there are others I'm missing. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | amelius a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Reading through this thread, I notice not many people are using models from the Chinese AI labs ... | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | donatj a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I feel like the odd man out ITT but Terra has been absolutely amazing for me and totally blows Sonnet out of the water. My company pays for Claude so I use Sonnet in the office but at home I use Terra and I find it far more likely to one shot some very decent code. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | NishanStepak a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I use Replit, but am looking for something less expensive, possibly open source. I see some very powerful open source systems that are cheaper, but I don't understand them that well. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | herpdyderp a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
- Default: Codex with Terra Max (because it's crazy cheap) - Preferred: Claude Code with Opus 5 Medium | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | rpmisms a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Gemini is the least stupid/most neurotypical model family IMHO. Works great for me | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | philbo a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Glm-5.3-flash for me. I like faster models because I stay involved all the way through. I don't delegate full control to the agent because it's harder to understand the end result that way. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | shamsalom94 21 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Claude, even though it constantly writes too much waffle and is generally quite slow. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | ewindisch a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I'm using a lot of gpt-5.6-Luna and glm-5.3-flash. Astra is really fantastic but it's too expensive. I average about 50-90B/tok/mo. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | adar2378 a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
So for Opus... but trying out astra occationally as well, might join openai boat if it keeps on serving good results | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | ghosty141 a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Terra Medium/High for most implementations, Sol High for tracking down hard bugs and planning more complex systems. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | nickthegreek a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
luna (max reasoning). price, as it sips quota on the Plus plan and is pretty competent for my needs. else I pretty much jump to Sol (high). i find that luna can work through a 5hr quota window on a /goal without hitting 0%. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | relug a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
i use opus5 and when i run out of usage i swap api to deepseek flash 4.1, deepseek better than opus now i think | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jamesponddotco a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Given that Fable is the only one that fucking listens to me and writes proper Go code, that’s what I use for everything code-related. At least on my personal projects. If I don’t care about the code, i.e. I’m writing something quick just to test something, I’ll use GPT Astra to save my Fable tokens. For conversations and research I usually go with GPT Astra Pro. For Home Assistant I use a combination of Grok 4.20 and DeepSeek Flash 4.1. Grok as a voice assistant, because it’s the only model with good latency here in Brazil (it’s nearly instant), and DeepSeek for everything else. I don’t really use LLMs for writing, but I’m writing a fiction book on the side, and when I tried to use one for ideas, Gemini Flash 3.8 gave me the closest to good writing out of the ones I tested. Not enough for me to use it for the task, though. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | JodieBenitez a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Whatever is the default in Codex CLI. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | guilhas 15 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I use several models, each seems to solve different problems better | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Mohamed_Amineio a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
opus actually strong at anything but it wastes token like drinking milk | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | w22oop a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I use gemma 4 12B and Qwen 3.8 9B | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | VariousPrograms a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GLM 5.3 Flash for personal hobby stuff. It's the first local-ish model that feels capable enough to be a default model to me (and it's very cheap). I'd rather get addicted to an open weight LLM in a class I can theoretically run if consumers ever get access to RAM again and can't get taken away from me. Locally I use Deepseek Flash, Qwen3.8-Next-Flash, or Gemma 4 depending on how slowly I want my slop to generate. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | behole a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Glm 5.3 Flash and DeepSeek 4.1 Flash | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Cakez0r a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I've been trying Grok 4.6 (High) alongside Claude Opus 5 (High). Grok is capable, but disappointing compared to Opus. My experience with Grok has also made me a little skeptical of benchmarks, because the two models benchmark similarly but it (subjectively) feels like there is quite a big capability gap between the two. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jraedisch a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
opusplan, for the same reasons. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Apreche a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I don’t use any LLM at any time for any reason. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | sidibe a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Gemini 3.8 + antigravity. Opus also in antigravity for harder stuff. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | cyanydeez a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Qwen3.8 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Havoc a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
GLM5.3 - I'm on one of the ancient plans, meaning basically unlimited. ...and then sprinkle in some other models when i think a second opinion will help | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | bellowsgulch a day ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
mimo-v2.5-free, mimo-v2.5, deepseek-flash in that order, honestly don’t even bother using qwen3.6-35b-a3b now unless i need uncensored tasks finished, mostly reverse engineering most engineering tasks don’t require frontier llms when they get stuck, then i consider moving up to more capable models purchasing a claude plan seems widely unnecessary to me the tasks they do better than the average engineer cut both ways: unless you have an existing portfolio of well written and designed work done pre-llms, it looks like you’re producing slop that pretends to be well designed poor typography choices despite using the mode, poor layout choices despite using popular CSS frameworks etc bad engineers will always be bad engineers tools don’t make up for it edit: a follow up to this— everyone is using eyebrows in their layouts and have no fucking clue why it was done to begin with everyone has a status pill floating above their front page hero display text and its not fucking status related so gross | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||