Remix.run Logo
joshka 4 hours ago

> It can't magically know what you want to say

I think for this argument to be true, the axiom that supports it is that the models have just as much context as they will ever have, and you cannot see being able to give them more / enough to be able to understand your perspective. That feels unlikely to be a position that doesn't change. As a society we're giving more and more context each day to this, and that makes this a valid opinion now, but one that erodes over time.

slopinthebag 2 hours ago | parent | next [-]

its not about dumping more and more info into the context, its about the intention. whats not in the context is just as important as what is. and i dont see how that can be automated.

also were seeing models become worse at writing as they get smarter.

owebmaster 4 hours ago | parent | prev [-]

Context went from 8,192 tokens on GPT 4 to 1M tokens currently with zero improvement. The latest models got even worse.

joshka 2 hours ago | parent [-]

Size of context is not the entire story here, it's ability to properly feed and index the context that's needed on this sort of thing. E.g. your entire slack/discord/email/github/jira/zoom meeting/coffee chat ... history is the context that you bring to the table on this sort of thing. Most of this is unindexed. Much of this will not be in the future.

> The latest models got even worse.

Which models? This is one of those things that likely has both model and domain specific aspects that impact your experience. In my experience with OpenaAI models predominantly (I previously worked there), they've improved significantly over the last 6-12 months. My experience with Claude is worse, but I haven't spent as much time getting into a mechanical sympathy there. They're still not perfect though and I have many steering docs that help avoid the biggest problems in the models I use when generating docs.

flipthefrog an hour ago | parent | next [-]

Claude writing quality got unbelievably bad with Opus 4.7, with no improvement in Fable. Opus 4.6 was fine. Im starting to see it as a security risk - my brain just can't process its word vomit, so just tell it to go on, implement whatever

2 hours ago | parent | prev [-]
[deleted]