Remix.run Logo
andy99 2 hours ago

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.

afavour an hour ago | parent | next [-]

What are you asking that you’re so regularly running into censorship?

alain_gilbert 7 minutes ago | parent | next [-]

The other day, I told claude that my physical wifi door unlock push buttons is a security risk because someone could run away with it and then unlock the door from outside whenever he wants. Then I told it that I want to introduce a concept of public/private key to uniquely identify my push buttons so that I can disable them individually using some crypto like ed25519...

Fable understood it as something along the lines of:

"introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED

gck1 17 minutes ago | parent | prev | next [-]

Reverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed.

It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.

Levitz 3 minutes ago | parent | prev | next [-]

I routinely get into blocks when running medicine-related material through it.

wild_egg an hour ago | parent | prev | next [-]

I'm doing a bunch of x86_64 assembly these days and Fable is simply not allowed to debug it. Hoping Opus 5 has a bit more freedom.

Retr0id an hour ago | parent [-]

I haven't been using it for long, but so far the refusals seem about on par with how things were on Opus 4.8.

wewtyflakes an hour ago | parent | prev | next [-]

I've hit it with intensely benign things; like asking it to make me a web-based client-side word game. I am guessing it saw the dictionary and pattern matched on various words, though ultimately it provided no explanation for why it triggered safeguards.

patcon 23 minutes ago | parent | prev | next [-]

Working on dimensionl reduction algorithms, I hit it all the time. I'm also trying to port related protocols from single-cell transcriptomics to collective intelligence systems (working with people x reaction matrices as analogous to single-cells cell x gene matrices.

Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me

weird-eye-issue 19 minutes ago | parent | prev | next [-]

Literally anything related to nutrition, athletic performance, etc especially if you ask it for research or sources

AnotherGoodName 11 minutes ago | parent [-]

Writing an implementation of a board game and one of the cards is called "microbes". Instantly knocked down to a lower tier model whenever it encounters that keyword because clearly bioweapons. Sigh.

msp26 an hour ago | parent | prev | next [-]

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

skinfaxi 6 minutes ago | parent | next [-]

Wait wtf. The mitochondria thing is true.

> Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats.

> Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.

Mistletoe 2 minutes ago | parent | prev [-]

Gemini knocks this out of the park, Gemini gang unite.

https://share.gemini.google/34vZzlnsmTaL

icedrift an hour ago | parent | prev | next [-]

If you even broach language related to biology you’ll get rerouted. I was presenting data in a grid and referred to a grid cell, Fable saw the word “cell” and safeguards kicked in

jefftk an hour ago | parent [-]

I thought we were talking about Opus 5, the model Fable now falls back to?

arcanemachiner 27 minutes ago | parent | prev | next [-]

I was profiling a slow machine the other day, and triggered the safeguards.

I've been saying this a lot lately, but it doesn't bites you until it bites you.

The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.

thousand_nights an hour ago | parent | prev | next [-]

i do homebrewing and asked it to compare some beer yeasts for me and hit the safeguards because... biology i guess lol

cute_boi 25 minutes ago | parent | prev [-]

Just ask math question and it will censor that. Even Misanthrophic employee confirmed that.

gck1 23 minutes ago | parent | prev | next [-]

Every time an online chatter (e.g. "limits are better", "model is better") makes me to reevaluate my principle of never paying Anthropic, I go to the model card, which strengthens my belief in the principle.

Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?

kccqzy 13 minutes ago | parent | next [-]

Just go to /config. The very second configuration item is “Switch models when a message is flagged” and presumably you want to turn this off.

Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?

markasoftware 17 minutes ago | parent | prev [-]

I don't believe they do silent downgrades right now, they're loud about it.

buzzerbetrayed an hour ago | parent | prev | next [-]

Yep. I cancelled my Claude Max subscription 2 weeks ago after feeling like Anthropic was doing everything it could to fuck with my day to day. Their lead would have to become significant for me to ever go back.

pinkyboy an hour ago | parent | prev [-]

[flagged]