| ▲ | Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra(cognition.com) |
| 211 points by seelos 4 hours ago | 92 comments |
| |
|
| ▲ | postalcoder 3 hours ago | parent | next [-] |
| If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?" |
| |
| ▲ | mediaman 3 hours ago | parent | next [-] | | This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on. | | |
| ▲ | postalcoder 2 hours ago | parent | next [-] | | > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing. | |
| ▲ | willcmcc an hour ago | parent | prev | next [-] | | "Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed" Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound. | |
| ▲ | nrmitchi 2 hours ago | parent | prev | next [-] | | > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Yes. | |
| ▲ | dpweb 2 hours ago | parent | prev | next [-] | | Fixed benchmarks will be debunked eventually I'd think. Better to use synthetic problems. Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that. | |
| ▲ | iLoveOncall 3 hours ago | parent | prev | next [-] | | > Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Yes? Just like every single model from every single AI lab. | |
| ▲ | felixgallo 3 hours ago | parent | prev | next [-] | | Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with? | | |
| ▲ | letmevoteplease 3 hours ago | parent [-] | | > Altman was caught in previous attempts trying to game benchmarks Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases. > does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with? I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities) And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves. | | |
| ▲ | kzrdude an hour ago | parent | next [-] | | There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback. That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user. Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think. | |
| ▲ | vlovich123 2 hours ago | parent | prev [-] | | Yeah and Astra is much better still |
|
| |
| ▲ | general_reveal 3 hours ago | parent | prev [-] | | I take it to mean the benchmarks are a marketing line item, as in, to sell this fucking thing you have to go out there and lie and the way everyone is lying is by doing exactly that, lying. They build for benchmarks and build benchmarks for builds. You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world. Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase. |
| |
| ▲ | fallingbananna 3 hours ago | parent | prev | next [-] | | Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8% | | | |
| ▲ | walrus01 2 hours ago | parent | prev | next [-] | | For comparison Qwen 3.8-Flash-Next which runs in under 190GB of RAM locally scores 25.3% on terminalbench 4.0. | |
| ▲ | throwatdem12311 2 hours ago | parent | prev | next [-] | | This is why I find benchmarks absolutely worthless. First, almost all models are within spitting distances of eachother. Second, it never translates to being better for my own workloads. You just need to make your own benchmarks. | |
| ▲ | Readerium 2 hours ago | parent | prev | next [-] | | Yeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0 | |
| ▲ | eranation 3 hours ago | parent | prev | next [-] | | When a benchmark becomes a target, it's no longer a good benchmark... | | |
| ▲ | tonychang430 an hour ago | parent [-] | | people are just fighting for numbers.. i don't fundamentally see the model being better |
| |
| ▲ | thefourthchime an hour ago | parent | prev | next [-] | | Came here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4. I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now! Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark. | |
| ▲ | enraged_camel 3 hours ago | parent | prev [-] | | Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far. | | |
| ▲ | nullbio 3 hours ago | parent [-] | | Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle. | | |
| ▲ | throwup238 3 hours ago | parent | next [-] | | > Andreessen Horowitz is being played like a fiddle. Andreessen Horowitz is not being played like a fiddle here. This might be their only investment in a decade that isn’t entirely predicated on being a scam. | |
| ▲ | thereitgoes456 3 hours ago | parent | prev [-] | | The Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less. While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others. Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work. | | |
| ▲ | selectodude 2 hours ago | parent [-] | | The cursor acquisition just shows that there’s always a dumber shithead out there. Though vaporizing Elon Musks money is about as pure of a good as there is out there these days. | | |
| ▲ | throwaway240403 2 hours ago | parent [-] | | Improvements in models and products coming out of SpaceXAI since the acquisition would seem to disagree with you. |
|
|
|
|
|
|
| ▲ | gruez 3 hours ago | parent | prev | next [-] |
| Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt. |
| |
| ▲ | fishtoaster 2 hours ago | parent | next [-] | | A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe? | | |
| ▲ | darkwizard42 32 minutes ago | parent | next [-] | | I think the evolution of the harness and ability to preserve loop context outside the context window has made running these kinds of agentic experiences easier. sorry so many buzzwords to say, the capabilities to do this kind of work are more accessible and easier to manage, so now it works! Good to see, and agree they were severely overhyping their product back then. | |
| ▲ | htrp an hour ago | parent | prev [-] | | a reminder that money does enable you to make mistakes and buys you the ability to correct from them | | |
| ▲ | unshavedyak an hour ago | parent [-] | | > and buys you the ability to correct from them Or at the very least, make more mistakes. |
|
| |
| ▲ | sterlind 7 minutes ago | parent | prev | next [-] | | I'm surprised they can post-train a closed model off of K3. Is that the norm for open-weight licenses? | |
| ▲ | notfromhere 3 hours ago | parent | prev | next [-] | | Well the models did get better but yeah their early product was godawful | |
| ▲ | esafak an hour ago | parent | prev [-] | | Previous versions were based on Kimi too. I'd consider if it I could access the model outside Devin. No lock in for me, thank you very much. |
|
|
| ▲ | nullbio 3 hours ago | parent | prev | next [-] |
| Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs. |
| |
| ▲ | notfromhere 3 hours ago | parent | next [-] | | This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI. Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own. Same reason Harvey is doing models now and basically every other provider | |
| ▲ | eru 3 hours ago | parent | prev | next [-] | | > If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else. | | |
| ▲ | bayesianbot 3 hours ago | parent | next [-] | | 1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead. btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I really didn't expect it to be anywhere near this good so we'll see where it ends up. And it's really fun throwing crazy amount of tokens at the wall for ~free instead of watching the subscription limits tick closer while your agents churn away. | |
| ▲ | colingauvin an hour ago | parent | prev [-] | | OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them. On the one hand you, if you bought a lot of compute a couple years ago (perceived demand, perceived shortage) you are in a good spot temporarily. But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be. I can almost, almost run DS4.1 Flash at home. 4 sparks can do it at 200+ tokens per second. I have two Sparks, so I am not in the club. Neither is your average laptop owner or gamer either. But your average HN software engineer can probably easily swing 2 sparks. |
| |
| ▲ | pizza234 3 hours ago | parent | prev [-] | | > we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). DS 4 Flash requires large amounts of memory to run at reasonable quants (I think a system with 160 GB or so). DS 4.1 Flash is even larger, I think around 250 GB. Any DS version is dumb when compared (in realworld tasks) to Astra/Opus 5, which means, one would spend thousands of dollars, and still need to rely on cloud services to do jobs that are non trivial. |
|
|
| ▲ | TheJCDenton 3 hours ago | parent | prev | next [-] |
| > SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice. |
| |
| ▲ | htrp an hour ago | parent [-] | | why are all American AI models basically Kimi in a trench coat | | |
| ▲ | ianm218 an hour ago | parent [-] | | No one in the US is going to fund pretty good open source with VC money. US has OpenAI/ Anthropic/ Google/ meta/ SpaceX atleast trying to make frontier foundation models 2 are using VC money + cash flow, last 3 are mainly cash flow + equity and debt. China has state banks and similar willing to fund lower margin open source labs. |
|
|
|
| ▲ | pkilgore 3 hours ago | parent | prev | next [-] |
| Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards. |
| |
| ▲ | AznHisoka 2 hours ago | parent [-] | | I have not met a single person/company that uses Devin… does anyone here actually use it? | | |
| ▲ | fschuett an hour ago | parent | next [-] | | I only used their "DeepWiki" automatic docs, they are pretty decent at getting an overview of a large project and are relatively accurate, with diagrams and anything. Haven't tried out their coding agent stuff. | |
| ▲ | PolCPP 2 hours ago | parent | prev | next [-] | | I pay for it (mostly because they grandfathered me from the old prices)! I used to use windsurf as my main editor until they changed their pricing model. Now i use it just to burn my weekly tokens on fable/astra if i remember to that on a task and that's it. | |
| ▲ | chris_st an hour ago | parent | prev | next [-] | | I use it, and have been happy with it for the most part. Like sibling, I use it for GPT-5.6-Sol and Opus work, and use their free models (GLM 5.2 for the past few months, trying SWE-2 now). | |
| ▲ | dominotw an hour ago | parent | prev [-] | | my friends at infosys are being trained on it |
|
|
|
| ▲ | andai 30 minutes ago | parent | prev | next [-] |
| Their benchmark used to show other metrics, like output tokens and time, but now only shows cost: https://cognition.com/frontiercode Which is too bad, since all of the gains here appear to be from massively reduced output tokens? The model SWE-2 is based on, Kimi K3, is cheaper per token than Sol, but costs more per task (ArtificialAnalysis) due to using way more tokens. Whereas, based on the graphs, SWE-2 appears even more token-efficient than Sol! That might have been worth showing off, if true. |
|
| ▲ | sbseitz an hour ago | parent | prev | next [-] |
| Why doesn't clickbait trash like this get moderated ? |
|
| ▲ | CyLith 2 hours ago | parent | prev | next [-] |
| I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actually need a jack-of-all-trades model to back my coding agents. |
| |
| ▲ | skrhee 2 hours ago | parent [-] | | I'm also in simulation software! Wondering which models you are finding helpful, the models I'm using for general SWE skills are horrible at our simulations and even basic physics/engineering calculation and intuition |
|
|
| ▲ | mydreamof 4 hours ago | parent | prev | next [-] |
| Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks? |
| |
| ▲ | harmonic18374 4 hours ago | parent [-] | | Probably, FrontierCode is made by Cognition itself. The model also seems worse in every way than DeepSeek v4.1 Flash, launched today. Also the submitter's account is very new which makes me suspicious of self-promotion. |
|
|
| ▲ | bobtheborg 3 hours ago | parent | prev | next [-] |
| SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it. Looking forward to 2 -- maybe it'll be usable |
|
| ▲ | Take8435 2 hours ago | parent | prev | next [-] |
| Post made by account 2 days ago. |
|
| ▲ | pelorat 39 minutes ago | parent | prev | next [-] |
| Unless it can do CAD via computer-use how can you say it rivals GPT-Astra? |
|
| ▲ | alansaber an hour ago | parent | prev | next [-] |
| Fair enough that they did a "propoganda and censorship" eval but not sure why i'd care about that in my highly juiced SWE kimi FT. |
|
| ▲ | monkeydust 4 hours ago | parent | prev | next [-] |
| As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft. There is an English translation somewhere. |
| |
| ▲ | ccapitalK 2 hours ago | parent | next [-] | | Pareto frontiers are pretty commonly invoked to describe tradeoffs in computer science and have been for quite a while. I remember the term being used in one of my early algorithms courses to describe the tradeoff between data structures with fast writes, ones with fast reads and ones that tried to balance the two. IIRC cognition boasted about hiring a lot of competitive programmers and algorithms experts back when they released Devin, so it tracks that they'd use the term. | | |
| ▲ | walrus01 2 hours ago | parent [-] | | It's interesting watching people throw about pareto frontiers sort of like how RF nerds approach the shannon limit (in a practical real world sense of the term, like charting possible modulations/data rates on a two way satellite modem's manufacturer datasheet). |
| |
| ▲ | arrowleaf 2 hours ago | parent | prev [-] | | Pareto leaked out of the sociology/econ bubble a long time ago :) Pareto principle, Pareto efficiency, Pareto distribution have been in the pop-sci buzzwords for quite awhile, I probably encountered it first in the 4-Hour Workweek. I don't think you can read a self-help book without the author introducing it as a groundbreaking principle to live your life by. |
|
|
| ▲ | eyeris 3 hours ago | parent | prev | next [-] |
| Wonder if this was the model that drove factoring the rsa-260 The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve. |
|
| ▲ | scronkfinkle 4 hours ago | parent | prev | next [-] |
| Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it. |
| |
| ▲ | samyok 3 hours ago | parent | next [-] | | SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI (https://docs.devin.ai/cli) :) Disclaimer: I work at Cognition, although was not involved in SWE-2 | | |
| ▲ | scronkfinkle 3 hours ago | parent | next [-] | | But I don't want to use your CLI. I already have my own harnesses and workflows. The friction is too high to "just try out" a new model like this. It would be preferable if I can evaluate it over, say, open router like all the other models and then decide from there if it's worth downloading a bespoke tool chain for only 1 lab's models | |
| ▲ | wren6991 2 hours ago | parent | prev | next [-] | | Your own CLI? Not even a /v1/chat/completions API? Is your business model based on pretending LLMs are not an interchangeable commodity already? | | |
| ▲ | anthonypasq an hour ago | parent [-] | | they are an agent company not a model provider, is this that difficult to comprehend? |
| |
| ▲ | randomblock1 2 hours ago | parent | prev | next [-] | | I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage. It does say 75% off though. Seems like for Pro subscribers SWE-1.7 is free, maybe SWE-2 is free for them? | |
| ▲ | vopi an hour ago | parent | prev [-] | | Heads up: it doesn't appear to be available on the Devin CLI (for me as a free user). |
| |
| ▲ | CamperBob2 3 hours ago | parent | prev [-] | | If its weights are open, that covers a multitude of other sins. Sufficiently-strong performance on the part of the new model would justify adapting existing tools to work with it. |
|
|
| ▲ | llmslave 3 hours ago | parent | prev | next [-] |
| At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!). I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around. Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job |
| |
| ▲ | hnedeotes 3 hours ago | parent [-] | | that's why after 1 year of product development of these AI 20x maxxed speed, we reached AGI 'wizards', there's really no difference in output, outstanding bugs no longer get solved and sites still suck, even doing things that were just regular development 20 years ago. Are you sure they aren't only producing 2.5% of your output that you manage just by farting into your phone? Are you sure it's 25% really? Seems way to high, days when I have diarrhoea my AI agents move even faster | | |
| ▲ | llmslave 3 hours ago | parent [-] | | please keep thinking this so i can relax with my automated job | | |
| ▲ | hnedeotes 2 hours ago | parent [-] | | It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu |
|
|
|
|
| ▲ | Tsarp 4 hours ago | parent | prev | next [-] |
| "SWE-2 is post-trained from Kimi K3" |
| |
| ▲ | airstrafer 3 hours ago | parent | next [-] | | Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model. Maybe still worth it if their "64% cheaper" figure holds. | | |
| ▲ | Bolwin 3 hours ago | parent | next [-] | | I don't think you know what distill means | | |
| ▲ | airstrafer 2 hours ago | parent [-] | | I guess I don't. Does post-training from another (larger) model not fall under the umbrella of distillation? I'd imagine it leads to the same spiky-ness issues...? | | |
| ▲ | FergusArgyll an hour ago | parent [-] | | Distilling you don't have the actual model weights of the teacher. All you have are the teachers answers to a lot of questions. You then teach your own smaller model to answer more similarly to the big teacher model. Fine tuning you have the actual model weights of the original model, you then train that model to answer in a different (or better) way. |
|
| |
| ▲ | Tsarp 3 hours ago | parent | prev [-] | | With the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2. | | |
| ▲ | samyok 3 hours ago | parent [-] | | SWE-2 is free for all subscribers on the CLI to try out for the next month :) | | |
|
| |
| ▲ | xlbuttplug2 3 hours ago | parent | prev [-] | | I presume post training is significantly easier than the distillation/training the top Chinese labs are doing. I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models. |
|
|
| ▲ | ltsSmitty 3 hours ago | parent | prev | next [-] |
| Well written and good diagrams. No idea the verity of the TMBB (trust me bro benchmarks) but it was pleasing to look at |
|
| ▲ | m3kw9 2 hours ago | parent | prev | next [-] |
| I'm using Codex, Gemini etc, they all have desktop apps and have a plan, how do i use SWE-2? Thats is a problem they have. I'm not about to switch out my workflow and plans with a shiny LLM that looks benchmaxxed and graph maxxed. |
|
| ▲ | wqash71 2 hours ago | parent | prev | next [-] |
| The horrible website is made by Claude or Cognition is distilled. I'm so tired of it all. |
|
| ▲ | _doctor_love 4 hours ago | parent | prev | next [-] |
| SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO. |
|
| ▲ | bluelightning2k 3 hours ago | parent | prev [-] |
| I like Cognition as a company and hope they succeed. Seemingly excellent engineering org. I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex. |