| ▲ | Muse Spark 1.3(developer.meta.com) |
| 266 points by bvaldivielso 3 hours ago | 162 comments |
| |
|
| ▲ | tyre 2 hours ago | parent | next [-] |
| Meta is one of those companies where, if there is anything remotely comparable, I'm happy to pay more to not use them. They've had a profoundly negative impact on society and Zuckerberg is not who I want controlling the future at the top of AI. I feel the same about Grok w/ Elon. I will pay extra to use someone else. I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money. And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom. |
| |
| ▲ | biddit 36 minutes ago | parent | next [-] | | Strong disagree with the Anthropic being good at all part. This is not defending anyone else, but… Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests. Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior. Aggressively scraping other people's works, despite the authors' requests not to do so. Then applying massive usage restrictions on their own work. And probably the most disqualifying is backing away from their own hard AI safety commitments. | | |
| ▲ | usef- 2 minutes ago | parent | next [-] | | Which safety commitments did they back away from? My understanding is that they believe safety can only be researched from the frontier, and so they're trying to be pragmatic to stay near the frontier (and viable) in their choices. From what I know, the "books3" dataset was normalised in the LLM and research ecosystem, where collected datasets were seen as valid to train on and/or fair use. I'm not sure any of the major frontier companies are free from that, if we don't believe it was fair use. I do think most of their choices are explainable by "they just believe in agi risk". You truly wouldn't want non-agi-pilled companies to train on your data and approach the frontier if you were worried. They are less worried about other "moral" decisions like "sharing" if they conflict with AGI. That definitely doesn't make them "good", but they do seem fairly "consistent". | |
| ▲ | dofm 19 minutes ago | parent | prev | next [-] | | > Anthropic leadership Which one? The main bit that reports to Daniela Amodei, or the little comfort blanket cabinet around Dario and his "chief of staff"? There is a leadership branch that can pretend to be morally qualified and aware and to think about the big picture and ethics. It is at least somewhat remote from the bit that is doing the actual business things. | |
| ▲ | nostromo 29 minutes ago | parent | prev | next [-] | | It makes me sad that people don’t see right through Anthropic’s gambit. They want to position AI as an insurmountable threat in order to regulate away any future competitors. They’re trying to speedrun regulatory capture. | | |
| ▲ | sscaryterry 25 minutes ago | parent [-] | | It is so obvious yet most people don't want to see it. | | |
| ▲ | metadat 23 minutes ago | parent [-] | | It's more like Anthropic present themselves in a deceptive way. I was confused at first too until someone on HN clued me in! Humans naturally want SOMEONE to be the good guy! Sad story, in this instance. |
|
| |
| ▲ | vovavili 8 minutes ago | parent | prev | next [-] | | It's almost like running a trillion-dollar business with neck-to-neck competition with other frontier labs and even state-sponsored efforts requires some ethical trade-off. | |
| ▲ | ACCount37 7 minutes ago | parent | prev [-] | | Pirating books is just straight up morally correct. I don't like Anthropic's bullshit "safety" filters, but training on shadow library data? Yeah no, it makes sense. It makes a lot more sense than having to work around copyright by scanning out physical books. Unfortunately, one was ruled legal and the other was not. |
| |
| ▲ | jwitthuhn an hour ago | parent | prev | next [-] | | Dario's idea of an ethical focus seems to be keeping powerful models out of the hand of anyone unethical, which coincidentally is everyone except him. | | |
| ▲ | tyre an hour ago | parent | next [-] | | Yeah this would be a great point if it were true and they didn’t give Mythos access to companies to fix bugs, which they did and have. It’s genuinely a difficult question. Not black and white. The models are really good at finding bugs, as demonstrated by people using Fable to reverse engineer. People make it sound like he’s just making it up. | | |
| ▲ | throwaway63486 17 minutes ago | parent | next [-] | | I'm the guy you replied to, apologies for using a different account I'm away from my computer now. The distinction to me is that Anthropic gives access to that model but doesn't give control. They reserve the right to cut you off if they don't like what you are doing and require you allow data retention for Fable and Mythos to ensure your are not up to any skullduggery. Meta, Alibaba, Mistral, even OpenAI has released models users can run locally and fully control. That is a whole world of difference. | |
| ▲ | hgoel 29 minutes ago | parent | prev | next [-] | | They gave access, but considering that they wouldn't even sign the "don't ban open weights" letter, it's clear they would prefer to have tight control over who they bless with that access. | |
| ▲ | phoghed 30 minutes ago | parent | prev [-] | | This would be more convincing if mythos was something uniquely special and not something merely a couple months ahead of everyone else. It was great marketing though. |
| |
| ▲ | porphyra 42 minutes ago | parent | prev | next [-] | | Dario's "ethical" look is also kinda sus. I hate to use ad hominem, but the dude's wife literally pitched a porn film to Epstein even after he was a convicted registered sex offender [1]. Dario is also really sinophobic (it is commonly claimed in Chinese AI circles that his former employment at Baidu triggered him so much that he harbors a personal grudge against the entire race). [1] https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-... | |
| ▲ | tehlike an hour ago | parent | prev [-] | | Pretty much. That's even worse imho. |
| |
| ▲ | platinumrad an hour ago | parent | prev | next [-] | | Funny. Dario seems like the biggest snake in the industry to me and has leaned the hardest into doom marketing out of all of the influential leaders. With Altman (or Google), it's a transaction, and that's something I can live with. | | |
| ▲ | tyre 43 minutes ago | parent [-] | | I just don’t see how people have looked at what has happened with Mythos and the deluge of fixes from companies, then come to this conclusion. He has a really hard job. He errs on the side of conservatism in releasing and then people get Really Mad. Safeguards on cybersecurity are not great for Anthropic revenue! As evidenced by people getting pissed, moving to Sol, and them having a smaller market for what Fable can do. It’s clearly bad for revenue and not great advertising to say, “you can’t use this but here is a nerfed version that will annoy you and not solve important problems.” | | |
| ▲ | adriand 4 minutes ago | parent | next [-] | | And he drew a red line wrt the Pentagon's use of Anthropic's models for autonomous weapons and surveillance of American citizens, and he stood by it, even when the government took steps to materially damage the company. This required true courage. Name me another CEO, of any major American company, that has demonstrated this much fortitude. | |
| ▲ | SwellJoe 23 minutes ago | parent | prev [-] | | Anthropic/Amodei have been the most alarmist about model safety, so multiple things can be true. A lot of tech companies avoided scrutiny by sending bribes to Trump (naked corruption is bad, I'd rather nobody do that), Anthropic didn't...so, combined with their fear-mongering about the danger of Mythos and open models (which seems aimed at regulatory capture) and the lack of bribes flowing to the Trump administration, they got stepped on by the federal government based on the excuse Anthropic provided. I dunno. Everybody seems to be playing pretty dirty. Some people have a much longer history of that, though. Obviously, Meta and Musk are outliers even in an industry full of problematic behavior. | | |
| ▲ | felixgallo 19 minutes ago | parent [-] | | I don’t see how you can look at what happened with hugging face and keep up the facade of anyone being alarmist or faking it. |
|
|
| |
| ▲ | badsectoracula 26 minutes ago | parent | prev | next [-] | | > he seems to have the most ethical focus He wants to build a tech-god kept in chains whose power he parcels out to the unwashed masses he deems worthy like some sort of high priest of intelligence. And that is being charitable and going by the interpretation that he actually believes what he says. | | |
| ▲ | im3w1l 17 minutes ago | parent [-] | | Well what do you want? Presenting clear, desirable, and achievable visions and trying to build consensus for how AI should develop is crucial at this point in time. |
| |
| ▲ | drob518 2 hours ago | parent | prev | next [-] | | Gotta be honest that I’m tired of the “I hate Zuck and Meta so much” comments every time Meta does anything. Ditto Elon/X. Fine, I get it. I don’t like Zuck either. But the post is about Muse Spark 1.3. What do you think about that? If you don’t like it because Meta made it, then maybe just don’t use it and stay silent. | | |
| ▲ | duplessitous 2 hours ago | parent | next [-] | | Technology doesn't just spring into being, there will always be comments on the organizations that developed it. If you don't like them or find them repetitive, it is far easier to collapse them and move on then bend a stranger to your will | |
| ▲ | canadaduane 2 hours ago | parent | prev | next [-] | | I get it, but the underlying problem is: we don't have a society-wide, effective solution to counterbalancing extractive systems. Lacking a reliable label, we have to constantly signal what's on the ingredients list. | | |
| ▲ | drob518 an hour ago | parent | next [-] | | Okay, but the comment I reacted to was not that. It was simply (paraphrasing) “I won’t use anything from Zuck/Meta.” If it had been, “Be careful because I have insider information that Zuck/Meta is using Muse Spark to do <insert-nefarious-thing-here>, and here’s my substantiation for that…” I’d be okay with it. That’s interesting information that moves a conversation forward. But it wasn’t. It was just content-free “I don’t like Zuck” nonsense. | |
| ▲ | luckylion an hour ago | parent | prev [-] | | Are not all corporations extractive by nature? Google clearly is. That's obviously not the issue with that -- you don't see those comments on Google's AI announcements. | | |
| ▲ | owebmaster an hour ago | parent [-] | | Yes you see. They get downvoted and flagged fast because there's a disproportionate amount of current and ex Google employees and stockholders around. |
|
| |
| ▲ | noduerme an hour ago | parent | prev | next [-] | | What I'm tired of is the top story (or five) on HN every day announcing Spark Opus Fable Grok Gemini v4.1i3-F. Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards? And look, part of my job is to use these things and part of my job is to pick EC2 servers, too. The front page of HN is increasingly resembling one of those endless AWS pricing lists. And yeah, I don't like any of the people or companies building LLMs either. At least the griping is somewhat interesting by comparison. The model isn't news. The news on Hacker News is that other professionals feel the same way. | | |
| ▲ | hadlock 36 minutes ago | parent [-] | | I think a lot of people are curious where the "knee" is on gains and productivity, particularly in the agentic space, which is where the real value is. A lot of us are being forced to shoe-horn this stuff into existing products, and knowing how much of the task the model can do now, vs having to build a complex custom harness, is valuable information to have. A year and a half ago it took our dev maybe six weeks of struggling with LangChain to approximate what Claude + MCP server can do today. The MCP server took us perhaps 2 days to build and 3 more to get it production ready. Today that MCP server gets 2-3 commits per month. I absolutely want to know when new models come out. As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off. |
| |
| ▲ | georgespencer 40 minutes ago | parent | prev | next [-] | | > If you don’t like it […], then maybe just […] stay silent. You might consider following your own advice. | |
| ▲ | monster_truck an hour ago | parent | prev | next [-] | | That's not how any of this works my man. Must be nice to think you live a life where neither has had a profound negative impact on your day to day | |
| ▲ | dangoljames an hour ago | parent | prev | next [-] | | yeah, we should all just stfu because one internet dude is tired of hearing it | |
| ▲ | jesse_dot_id an hour ago | parent | prev | next [-] | | Staying silent is unfortunately how fascism festers. | |
| ▲ | whateveracct an hour ago | parent | prev | next [-] | | Zuck's bad PR is to blame here. Not the commenters. He should fix that. Anduril makes this same complaint whenever their job posts get dumped on. Same idea. Fix your bad PR, buddies :) | |
| ▲ | idiotsecant 2 hours ago | parent | prev | next [-] | | Not liking something because the embodiment of corporate malfeasance is a rational way to decide what products to support. | |
| ▲ | troupo an hour ago | parent | prev [-] | | > But the post is about Muse Spark 1.3. What do you think about that? That: - like all models it was trained on stolen data - additionally it was trained on Facebook users who were all opted in to AI training with a convoluted 10+ step process to opt-out of > If you don’t like it because Meta made it, then maybe just don’t use it and stay silent. Why should anyone stay silent? |
| |
| ▲ | dimgl 2 minutes ago | parent | prev | next [-] | | I'll use both Muse Spark and Grok. | |
| ▲ | aftbit an hour ago | parent | prev | next [-] | | Okay but Muse Glimmer 30B is one of the best small open weight models today, and IMO the best from a US lab (only real comparison is Gemma4 dense right now). | | |
| ▲ | tyre 42 minutes ago | parent | next [-] | | Totally fine with open weight, since other people can provide it and Meta isn’t making money. I’d use an AWS-hosted version. | |
| ▲ | Bluestein an hour ago | parent | prev [-] | | I am finding Poolside's a decent model.- |
| |
| ▲ | devy an hour ago | parent | prev | next [-] | | > I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money. Amodei is NO Saint!!! He's the most savvy in drumming up the AI doomsday scenarios and haven't yet to apologized his failed forecast of Claude taking over 90% of the coding jobs. | | |
| ▲ | samtheprogram an hour ago | parent [-] | | He didn't say 90% of the coding jobs. He said LLMs would write 90% of the code. As in be LLM generated. |
| |
| ▲ | ls_stats an hour ago | parent | prev | next [-] | | Is it really hard to understand that there's no good guys? Amodei, Altman, Zuckerberg, Musk, etc. They all sound the same to me. | |
| ▲ | redox99 2 hours ago | parent | prev | next [-] | | Google, Zuck, Sama, Elon, Amodei (in no particular order). They all suck. Pick your poison. | | |
| ▲ | kenjackson 2 hours ago | parent [-] | | They don't all suck equally. Here's the order, from best to worst. Amodei Google SamA Zuck Elon | | |
| ▲ | redox99 2 hours ago | parent | next [-] | | You can ask 100 people and they'll all give you a different list. It's subjective. I think a less personal ranking would be, as a business owner, which of those providers is more dependable? As in, you don't care about evil, just your stuff working. I think maybe OpenAI? | | |
| ▲ | a2ff6eeb0 an hour ago | parent | next [-] | | Google. They have experience operating at scale, and AI is a big enough focus that they won't wind it down. All the big providers are kinda crappy, but if you want reliability, Google is the best option. | | |
| ▲ | redox99 an hour ago | parent | next [-] | | Definitely not Google, countless horror stories and infamous for killing stuff. OpenAI is still serving GPT 3.5 turbo as far as I remember. | |
| ▲ | ralusek an hour ago | parent | prev | next [-] | | Google, famous for not winding things down. | | | |
| ▲ | applfanboysbgon an hour ago | parent | prev [-] | | Google has experience working for themselves at scale. Your business should never rely on Google more than it is forced to. Even if it's not something they'll wind down, providing acceptable service to anyone is not on their agenda. GCP speaks for itself... |
| |
| ▲ | utopcell an hour ago | parent | prev [-] | | > You can ask 100 people and they'll all give you a different list. True. With 5 choices you need at least 126 people before you can guarantee that two lists are the same. |
| |
| ▲ | runarberg 3 minutes ago | parent | prev | next [-] | | They all suck beyond any tolerable threshold. Some of them are further away from the threshold. But at this point, how far each is from the tolerable threshold is besides any point and not worth arguing over. The least of five evils is still evil. | |
| ▲ | cactca an hour ago | parent | prev | next [-] | | Demis Hassabis is, by any standard, the most ethical of the bunch. | |
| ▲ | scottyah 2 hours ago | parent | prev | next [-] | | SamA better than Zuck? Zuck was at least a kid when he made a lot of his bad decisions, and he seems to be getting much better. Sam is on the reverse trajectory. | | |
| ▲ | KptMarchewa an hour ago | parent | next [-] | | I am 100% convinced Zuck is maybe better at masking now, but is exactly the same lizard who wrote >>> They "trust me"
>>> Dumb fucks | |
| ▲ | Zambyte an hour ago | parent | prev [-] | | Sam is on a delayed trajectory of power, but he surely was not great when he was young either. See: Aaron Swartz calling him a sociopath who could not be trusted, well over a decade ago. |
| |
| ▲ | porphyra an hour ago | parent | prev [-] | | In my personal opinion (this will be controversial and feel free to disagree): Elon is the best. * great contributions to many industries including spaceflight, electric cars, and self driving cars. It doesn't even matter if he is the technical mind behind these achievements or if he is just a buffoon that pretends to know the implementation details; the dude has a way of bringing together experts, having the overall vision, and managing them properly to ship amazing stuff. * sane and reasonable takes on AI/LLM stuff. I can't really argue with "pursuit of truth" as the guiding principle. Grok talks normally without "Claudlish", has a balanced score on political bias unlike other models, has a low hallucination rate, is the best at dealing with latest news (unlike ChatGPT that refuses to believe new developments and gaslights the user), and they "never silently downgrade intelligence or fall back to other models." In contrast, while Dario is doubtless a super smart pioneer in the AI space, his sanctimonious "We know what's good for you" attitude and extreme censorship is really offputting. The lengths to which he tries to ban or hamstring open models seems like an underhanded way to defeat competition. If he were to succeed, it would be a big setback to the thriving ecosystem of open models and hamper the development of the entire industry. |
|
| |
| ▲ | optimalsolver 2 hours ago | parent | prev | next [-] | | If it was up to Dario we'd all be banned from using open-weight models, and we'd have to be investigated for PRC connections before sending our allotted five API queries a week. | |
| ▲ | spiderfarmer 21 minutes ago | parent | prev | next [-] | | Same with Grok. | |
| ▲ | loeg 2 hours ago | parent | prev | next [-] | | "Avoid generic tangents" / "Please don't complain about tangential annoyances." | | |
| ▲ | _diyar 2 hours ago | parent | next [-] | | How is this a tangential annoyance or a generic tangent? > Meta announces they have a new model, demonstrating its capabilities. > Parent comment states „regardless of this model‘s specific capabilities, if I can avoid it I will.“ | | |
| ▲ | loeg 2 hours ago | parent [-] | | Grandparent comment has zero to do with the article. It's just GP generically bitching about Meta. (Your "quote" of the comment does not appear anywhere in the actual comment.) |
| |
| ▲ | reaperducer 2 hours ago | parent | prev [-] | | "Avoid generic tangents" / "Please don't complain about tangential annoyances." That's pretty much 90% of HN these days. Apple releases a new iPhone? Here comes the flood of decade-old complaints about long-discontinued Mac butterfly keyboards and walled gardens. Microsoft releases a new version of Windows? Here come the gripes about Azure. Google changes something in GMail? Play Store! It's like there's an army of bots out there determined to reduce the productivity of the Western tech bubble by diverting everyone into endless circular arguments about absolutely nothing of relevance to the topic at hand. |
| |
| ▲ | bradlys an hour ago | parent | prev | next [-] | | What is the point of this comment? | |
| ▲ | TacticalCoder an hour ago | parent | prev | next [-] | | Meta and Microsoft are two of the absolute worst evil companies on earth and Amodei is trying very hard to join them. These Effective Altruists are despicable people: a bunch of thieves working to line up their own pockets while posturing as a force of good. Remember that they schemed to not only present SBF as the 2nd coming of Christ (including in the NYT and in Forbes) but to also give him a voice after his scam had been uncovered. Thankfully, the judge didn't have any of this Effective Altruist bullshit. SBF invested 500 millions of misappropriated funds in his buddy from the EA movement's Anthropic company (and, thankfully, the judge forced those shares to be sold: so SBF didn't get to be a billionaire). You cannot hate enough people who say that harming others for the greater good is justified. Then of course, already mentioned in this thread, there's the whole Epstein/Amodei's "I'm in the porn business" wife connection (where you don't need to squint much to see young women abused). These kind of people are the absolute worst scum on this earth. | |
| ▲ | fouc 2 hours ago | parent | prev [-] | | How have they had a negative impact? How about google? | | |
|
|
| ▲ | simonw 3 hours ago | parent | prev | next [-] |
| llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...4.2266 cents, 38 seconds. For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat. UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... The most expensive was reasoning level xhigh - 7.5 cents, 1m34s. And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... |
| |
| ▲ | drusepth 2 hours ago | parent | next [-] | | Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality. | | |
| ▲ | porphyra 2 minutes ago | parent | next [-] | | The canonical view of a bicycle is facing right. Usually, people want to draw/photograph/depict the side of the bicycle with the running gear, which is on the right side of the frame for historical reasons. | |
| ▲ | vunderba 2 hours ago | parent | prev | next [-] | | The more generic your prompt, the more generic the response. It's a regression to the "mean" of the training data aka GIGO for AI. It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows. In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs. | | |
| ▲ | polyterative 2 hours ago | parent [-] | | as a kid I did them like this. nobody told me to do that. are we all so similar? | | |
| ▲ | srcreigh 2 hours ago | parent | next [-] | | The adults brainwashed us https://www.ikea.com/ca/en/p/barndroem-box-beige-70560615/ https://www.ikea.com/ca/en/p/vallaby-rug-green-10548216/ | |
| ▲ | collabs 2 hours ago | parent | prev [-] | | I sincerely believe I've never had a single original thought™ in my whole life. There is this scene in the HBO series Westworld where a "host" says some words in sequence which is shown on a display as she says it. Of course, even me thinking of this scene and connecting it to your comment was not original, someone else clearly had the same programming as me. A medium blog post says > Pair what with me?” — the moment Maeve (a humanoid android) uttered those words in Westworld (Season 1, Episode 6: “The Adversary”), something clicked. Not for the average viewer, but for me, a STEM educator and AI enthusiast who, just weeks earlier, had read Stephen Wolfram’s seminal essay, What Is ChatGPT Doing … and Why Does It Work? |
|
| |
| ▲ | BeetleB 27 minutes ago | parent | prev | next [-] | | Not when rendered via POV-Ray: https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt... I plan to update it with more pelicans from all the models released since. (Spoiler alert: They haven't improved much since then). | | |
| ▲ | xhrpost 14 minutes ago | parent [-] | | Wow, I actually had this exact idea. I was specifically curious as to how well a given LLM could understand a DSL that hasn't changed much in a couple decades and doesn't have nearly as many examples to learn from online. Seems like it did alright, all things considered. |
| |
| ▲ | simonw 2 hours ago | parent | prev | next [-] | | It's really interesting, isn't it? They almost always cycle from left to right - but I have had a few which cycle in the other direction. The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration. | | |
| ▲ | m12k 2 hours ago | parent | next [-] | | It's my impression that it's common in western culture, where text is read left to right, and timelines are visualized as going from left to right, to also animate things going from left to right, since westerners thus have an instinct that "right = forward", so it "feels right" (familiar). I wonder to which degree this is reflected in the training data? And if you'd be more likely to get left-facing pelicans if you prompted it in Hebrew, Arabic or another right-to-left language? | |
| ▲ | johntb86 2 hours ago | parent | prev | next [-] | | Someone studied this (among other thigns): https://dylancastillo.co/posts/pelicanmaxxing.html . Pelicans on bikes always face right in this test, but other animals on other transportation methods sometimes face left. | |
| ▲ | piker 2 hours ago | parent | prev [-] | | I was going to ask the exact same question earlier but deleted it after thinking “I’m sure Simon has done some sort of discussion on this.” Since it does seem novel to you, too, it would be really interesting to read more about this phenomenon. |
| |
| ▲ | postalcoder 2 hours ago | parent | prev | next [-] | | Search Google Images for "bicycle". Almost all bicycle product shots are staged the same way: side view, going left-to-right. It makes sense to me that given that skew in the training data, the model grounds itself in the bicycle. | | |
| ▲ | daemonologist 2 hours ago | parent | next [-] | | and furthermore, this is because the drivetrain is ~always on the right side of the bike - if you want to inspect or admire a bicycle you look at the right side, as you might look under the hood of a car. (Why the drivetrain is on the right, I don't know. But most bike parts follow open standards so it's quite entrenched.) | | |
| ▲ | kibae 2 hours ago | parent | next [-] | | Since most languages read from left to right, rightward movement tends to read as forward progression. So when showing a bicycle in side profile, having it face right feels more naturally like it’s moving forward. | |
| ▲ | georgemcbay 2 hours ago | parent | prev [-] | | > and furthermore, this is because the drivetrain is ~always on the right side of the bike While I'm sure this factors into things for advertisements for bike components, there is also just a general preference that westerners have for left-to-right motion. Not just in bike ads, but all ads with (or suggesting) movement. And also not just ads, but movies where directors believe left-to-right motion is associated with progression and right-to-left motion is regressive. |
| |
| ▲ | threetonesun 2 hours ago | parent | prev [-] | | Product shots yes, people riding them its more like 50/50. Also if you search for a specific bicycle race you'll find more going right to left. |
| |
| ▲ | optimalsolver 2 hours ago | parent | prev | next [-] | | Sun is missing a few rays and not wearing sunglasses. | |
| ▲ | ModernMech 2 hours ago | parent | prev | next [-] | | Yes, I do a thing where I ask the machine to generate responses in the form of a lizard talking to a cat. The lizard is always a green gecko and the cat is always orange, which I never specify. | |
| ▲ | reaperducer 2 hours ago | parent | prev [-] | | Is there a reason these pelicans always have roughly the same composition Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do. Much like a mother pelican, they regurgitate what they've been fed. |
| |
| ▲ | hollowturtle an hour ago | parent | prev | next [-] | | Is there any point anymore regarding this svg test? I would not be surprised if in the training they're fine tuned for this task too | | | |
| ▲ | ipsum2 18 minutes ago | parent | prev | next [-] | | All of the links show "Error: Gist API returned 403". | |
| ▲ | tintor an hour ago | parent | prev | next [-] | | Did any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans. Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability. | | | |
| ▲ | jonahx 2 hours ago | parent | prev | next [-] | | If you have a grading rubric, huge points off for adding arms instead of using the wings as arms! | | |
| ▲ | Fergusonb 2 hours ago | parent [-] | | I think it's hilarious that this detail is enough for me to dismiss looking into the model, but here we are, and it is. |
| |
| ▲ | drob518 an hour ago | parent | prev | next [-] | | Simon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time. | |
| ▲ | jonplackett 2 hours ago | parent | prev | next [-] | | Has any ab tried to game this yet and just made the most amazing pelican by hand and always reply with that? | |
| ▲ | jttnr 2 hours ago | parent | prev | next [-] | | I wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans? | |
| ▲ | EugeneOZ 2 hours ago | parent | prev | next [-] | | Absolutely BRUTAL! :) Thank you for doing this, I love your benchmark the most! | |
| ▲ | tomrod 2 hours ago | parent | prev | next [-] | | What does the mean pelican look like at this point? Also 3X token use vs. 1.2 | | | |
| ▲ | 0xbadcafebee 44 minutes ago | parent | prev | next [-] | | For all the comments of "I'm sure they're fine-tuning for pelicans": https://dylancastillo.co/posts/pelicanmaxxing.html "Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing.
Or at least they’re not doing it in a plainly obvious manner."
| |
| ▲ | jmkni 2 hours ago | parent | prev [-] | | lol Definitely an upgrade over 1.2 |
|
|
| ▲ | superfrank 2 hours ago | parent | prev | next [-] |
| I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it. I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be. I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be. Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake. |
| |
| ▲ | sejje an hour ago | parent | next [-] | | If it's a mistake, it should course-correct. I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed. I have to think that's the future, somehow, and I'm really excited about it. | |
| ▲ | MangoCoffee an hour ago | parent | prev [-] | | >I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code. | | |
| ▲ | dakolli a minute ago | parent | next [-] | | Every lab trains their models with AI generated code at this point. | |
| ▲ | KptMarchewa an hour ago | parent | prev | next [-] | | I would imagine your interactions with it are more important than the output. | |
| ▲ | TiredOfLife an hour ago | parent | prev [-] | | Training on ai generated content is how the models got a big jump in capability |
|
|
|
| ▲ | Lucasoato 2 hours ago | parent | prev | next [-] |
| A model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim). Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction. |
| |
| ▲ | dbbk an hour ago | parent [-] | | How is it not SOTA? It's beating 5.6 Sol. | | |
| ▲ | ctolsen 43 minutes ago | parent [-] | | You gotta keep up. Fable 5.1 came out yesterday and is better so anything else is to be treated as garbage now. |
|
|
|
| ▲ | bertili 2 hours ago | parent | prev | next [-] |
| DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap!
Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down! |
| |
| ▲ | WASDx 2 hours ago | parent | next [-] | | With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace. | |
| ▲ | cbg0 2 hours ago | parent | prev | next [-] | | But is the score really reflective of the quality or are both models benchmaxxing? | | |
| ▲ | bermudi an hour ago | parent [-] | | Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD. This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol. |
| |
| ▲ | dominotw 2 hours ago | parent | prev [-] | | how much of it is from reallocation of staff to ai training and labeling |
|
|
| ▲ | jmward01 11 minutes ago | parent | prev | next [-] |
| muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a setting somewhere. This is the first quantifiable number I have seen out there from a model provider. Maybe it can help in lawsuits to quantify the damages for copyright/other things? |
|
| ▲ | apodolny 37 minutes ago | parent | prev | next [-] |
| I like the approach of providing a discounted version of the API that is used to train vs. the full price version. Seems reasonable and transparent. |
|
| ▲ | gehsty 27 minutes ago | parent | prev | next [-] |
| As a product, would developers switch to a meta model/harness? I don’t think so. Only way I see is if it becomes the new SOTA / frontier, does anyone think Meta will surpass Anthropic or OpenAI? I still can’t get my head around why language models are an existential threat to Meta - they own the platforms people watch adds on? |
| |
| ▲ | phyrex 2 minutes ago | parent [-] | | Meta also has 50k engineers. Not to mention that tons of meta infrastructure - including ads! - use AI. Would you want that sort of business be this dependent on someone else? |
|
|
| ▲ | jumploops 2 hours ago | parent | prev | next [-] |
| The "contributor" pricing is the standout here at a ~20x discount, if you allow training on your data. The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $). Stats: 1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok) |
| |
| ▲ | 2001zhaozhao 2 hours ago | parent [-] | | I have a feeling that Meta is not gonna like what people actually use the contributor model for lol. (It's probably going to be a bunch of repetitive batch jobs like web search that have no training value) | | |
| ▲ | hadlock 20 minutes ago | parent | next [-] | | There's a lot of value in agentic loop tool failure + recovery training data | |
| ▲ | winstonp an hour ago | parent | prev | next [-] | | It's the perfect model for open-source work because it's gonna end up in the training data anyway | |
| ▲ | dbbk an hour ago | parent | prev [-] | | Web Search doesn't have a discount on contributor pricing |
|
|
|
| ▲ | majerep 2 hours ago | parent | prev | next [-] |
| The previous version was, in my experience, the best free model available on OpenCode. It's been very good at simple/moderate tasks where I am precise in my ask and it doesn't need to make a ton of undefined assumptions. Hopefully this new version is also available on opencode for free. |
|
| ▲ | 7734128 2 hours ago | parent | prev | next [-] |
| Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists. |
| |
| ▲ | 0xbadcafebee 36 minutes ago | parent [-] | | I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models |
|
|
| ▲ | cnxhk an hour ago | parent | prev | next [-] |
| artificial analysis results: https://x.com/ArtificialAnlys/status/2095247787277553929 |
|
| ▲ | wxw 2 hours ago | parent | prev | next [-] |
| “contributor” pricing at $0.10/$0.20 is crazy cheap if it’s measuring up to Sol. Definitely shows how important a user data flywheel is for RL and model improvement. |
|
| ▲ | finnjohnsen2 2 hours ago | parent | prev | next [-] |
| So one model is "Not used to improve our products" and is 10-20 times more expensive to the "Used to improve our products"-model. Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room. |
| |
| ▲ | duplessitous an hour ago | parent | next [-] | | What is the confusion? They directly state that you are the product if you use their discounted offering. It isn't an assumption that should lead you to this, it is Meta's very direct communication that should lead you to this | |
| ▲ | Jcampuzano2 2 hours ago | parent | prev | next [-] | | I'm confused what your surprise is here. It's plain and simple right to the point wording. I don't see the wiggle room at all. | |
| ▲ | thefreeman 2 hours ago | parent | prev | next [-] | | aren't they explicitly saying this with both their pricing and their wording? I'm not sure what you are alluding to? | |
| ▲ | warkdarrior 22 minutes ago | parent | prev | next [-] | | Privacy is not free. They make it quite clear that they charge more if you don't want your data used by Meta. | |
| ▲ | whimsicalism 2 hours ago | parent | prev | next [-] | | the meaning is pretty obvious - they want to train on your chats & tasks and are willing to subsidize for the privilege of doing so. | |
| ▲ | zhoBEENG an hour ago | parent | prev | next [-] | | Would it help you understand if they were labelled "For Dumb Fucks" and "For Everyone Else"? | |
| ▲ | IshKebab an hour ago | parent | prev | next [-] | | I think it's more that the "not used to improve our models" is expensive because companies need that. It's simple price differentiation. In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that. | |
| ▲ | bigyabai 2 hours ago | parent | prev [-] | | Given OpenAI and Anthropic's behavior, do you really expect them to be singled out for this practice? Zero trust has been in "LGTM" territory for years now. Meta's bet against people taking a principled stance arguably paid off great. |
|
|
| ▲ | maciejgryka an hour ago | parent | prev | next [-] |
| Does anyone know what the license for this model is? Specifically any word on restrictions about what it can be used for? |
|
| ▲ | Gecko4072 2 hours ago | parent | prev | next [-] |
| Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages. |
| |
| ▲ | WASDx 2 hours ago | parent [-] | | I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks. |
|
|
| ▲ | souvlakee 2 hours ago | parent | prev | next [-] |
| Why they didn't use LLM to create html table instead of https://lookaside.fbsbx.com/elementpath/media/?media_id=1048...? |
|
| ▲ | LZ_Khan an hour ago | parent | prev | next [-] |
| Ha, even with monitoring engineers keystrokes and mouse movements not SotA on OSWorld. |
|
| ▲ | fibonacci112358 2 hours ago | parent | prev | next [-] |
| Is everyone rushing to launch something before Astra tomorrow? |
|
| ▲ | mromanuk 2 hours ago | parent | prev | next [-] |
| I didn't like 1.2, It make some mistakes in a web app, so I quickly went back to Claude, Kimi K3 or Deepseek V4. Hope this one can clear agentic development, because Muse Spark models are fast and cheap. |
|
| ▲ | geooff_ 2 hours ago | parent | prev | next [-] |
| Could this be best intelligence / $ if you're willing to let zuck digest your data? |
|
| ▲ | meerita 2 hours ago | parent | prev | next [-] |
| I declined the use of cookies and everything went black. No content at all. Dissapointed. |
|
| ▲ | improgrammer007 17 minutes ago | parent | prev | next [-] |
| All people here care about is hating Meta. Just look at the top voted comment. No one cares about the merits of the model, etc. HN has become nothing but an echo chamber. |
|
| ▲ | lostmsu an hour ago | parent | prev | next [-] |
| What a day. OpenAI is behind basically all major competitors - at least for a some amount of time. |
|
| ▲ | ChrisArchitect 2 hours ago | parent | prev | next [-] |
| Blog post: https://research.meta.ai/blog/introducing-muse-spark-1-3 (https://news.ycombinator.com/item?id=49541149) |
| |
| ▲ | sunaookami 2 hours ago | parent [-] | | >Previously available reasoning modes are available today with max reasoning coming shortly after we finish additional safety testing Lmao. And their benchmark table only shows max reasoning. |
|
|
| ▲ | tinyhouse 2 hours ago | parent | prev | next [-] |
| I had no idea Meta has a coding agent harness. Does anyone have experience with it and can comment? The 1.3 contributor prices look very attractive. I'll probably start using their API if performance is good and the API is reliable with decent rate limits. |
| |
| ▲ | meric_ an hour ago | parent [-] | | You should use their harness. They trained it on multiple harnesses but have specifically optimized it for their harness. Cline also did an independent experiment w spark 1.2 where using the native harness makes it use fewer tokens / turns to accomplish tasks |
|
|
| ▲ | frozenseven 3 hours ago | parent | prev | next [-] |
| This should probably be primary: https://news.ycombinator.com/item?id=49541149 |
|
| ▲ | IshKebab an hour ago | parent | prev | next [-] |
| Lol "not used to improve our models" is AI's enterprise SSO. |
|
| ▲ | dangoljames an hour ago | parent | prev [-] |
| If it's from meta, pit h in the bin. |