| |
| ▲ | NitpickLawyer a day ago | parent | next [-] | | It really isn't and it's sad seeing so many people say it so confidently on this site. It only detects plain / basic prompted stuff. "write me an essay on x", sure. The moment you prompt it differently, it stops working. Add "output should be STE100 compliant" or something to your prompts and detection goes away. Fine-tune any local llm on real human prose, and detection goes away. Edit 2-3 characters (emdashes, lists, etc) and detection goes away. And that's for just basic detection. There have been plenty of examples of 100% human written content (either old, or unpublished) that gets falsely flagged as AI. And these are just technical aspects. The main issue is that pangram and other solutions are being used to summarily judge students work, and that is orders of magnitude more fucked up. Accusing someone of cheating can have devastating effects on their education/career/etc. and they're doing it with snake-oil closed boxes, at scale. We really really shouldn't support this, especially here on a technical site. | | | |
| ▲ | jedberg a day ago | parent | prev [-] | | No it isn't. I've put human written material in there (my own unpublished material) and it says 90%+ AI. Then I put some AI material in and it said 30% chance of AI. I only tested it with the two pieces, but was not impressed. | | |
| ▲ | calmoo 21 hours ago | parent | next [-] | | I’ve never actually seen someone give an example of what they say pangram get wrong, so please go ahead and share. | | |
| ▲ | jedberg 19 hours ago | parent | next [-] | | You can read it here: https://gist.github.com/jedberg/e24124e577c0da1c8445c22798eb... I posted here on HN but it got removed. | | |
| ▲ | calmoo 16 hours ago | parent [-] | | I'll take your word that it's human written, but reading that gist, it absolutely reads like Claudeslop / GPT slop, it doesn't surprise me in the slightest that Pangram flagged it as AI when it reads identically to LLM output - this feels like a very acceptable edge case to me (assuming you are telling the truth). Are you absolutely sure you wrote this by hand? If so it's kind of remarkable how close to an LLM you write like. | | |
| ▲ | jedberg 14 hours ago | parent [-] | | It’s important to remember that LLMs were trained on well written human text. People who write well are going to sound like an LLM. Especially if it’s a marketing message for a website. I’ve been accused of being an LLM multiple times here on HN too. I know you have no way to know for sure other than trusting that I’m not using an LLM to write. But it’s pretty frustrating that people jump right to LLM accusations. | | |
| ▲ | calmoo 14 hours ago | parent | next [-] | | I’ll be honest, i’ve never read text that is human written that reads as close to an LLM as your sample sounds. I think discounting Pangram’s accuracy based on that sample isn’t very reasonable. Really nobody writes like that other than LLMs! | | |
| ▲ | jedberg 12 hours ago | parent [-] | | Here is another example: https://www.reddit.com/r/ExperiencedDevs/comments/1pyjkuf/i_... Today Pangram says it is 100% human, which is correct. But yet I got multiple DMs when I posted it 8 months ago saying "stop posting AI slop!" in response to that comment. At the time, Pangram marked it as 50% AI. So to Pangram's credit, they got better. | | |
| ▲ | calmoo 6 hours ago | parent [-] | | That comment doesn’t read like slop in the slightest to me, so the people DMing you have a bad eye for it. Regardless, I think your ‘human’ sample is not a good indicator of the quality of pangram. I would update your priors a bit. |
|
| |
| ▲ | ekelsen 11 hours ago | parent | prev [-] | | Do you have any published dateable text from pre 2023? I'm super curious if you always wrote like this, because it is exactly in the style of AI slop. People say it's in the training data, but I haven't seen any good examples of clearly dated pre 2023 text that sounds like this. | | |
| ▲ | jedberg 11 hours ago | parent [-] | | I have 1000s of reddit and hacker news comments from before 2023, and lots of long form writing too. But as you point out, those are all in the training set, and in pangram's "definitely human" training set too. The ones that sound like LLMs tend to be the ones that were well researched and spent more time on, not the off the cuff stuff, which is most of what I write. So it would take me a while to find something like that. But you're welcome to dive into my reddit and HN history, or all my blog posts on the wayback machine if you want to look for one. :) |
|
|
|
| |
| ▲ | 21 hours ago | parent | prev [-] | | [deleted] |
| |
| ▲ | ekelsen a day ago | parent | prev | next [-] | | Can you share your human piece that got flagged? And the AI piece that didn't? (Or snippets?) I've found pangram to be very accurate, so I'm curious. | | | |
| ▲ | matsemann a day ago | parent | prev [-] | | Everything has false positives or negatives. 2 examples doesn't tell anything about how good or bad it is. | | |
| ▲ | pixl97 19 hours ago | parent [-] | | Being that Panagram lies about its effectiveness in their presentations while hiding it's actual capabilities and testing methods rather deeply I have little faith in it. (1 in 10000 wrong in presentations verses 2-4% wrong in testing). For example they have a corpus of older pre-llm text and use that as the example their current model doesn't misclassify human written text. It shouldn't take much thinking to realize why this is a fucking stupid benchmark. Every day humans use LLMs and read LLM content Panagram becomes more useless because it forces languages to have a stopping point sometime around 2020. If you adopt any LLMism or are one of those unlucky people that already talked like an LLM before LLMs then all your shit is getting marked even though it was created by the human mind and written by human hands. |
|
|
|