| ▲ | The Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news(lenfestinstitute.org) |
| 73 points by giuliomagnifico 3 days ago | 26 comments |
| |
|
| ▲ | syntaxing 3 days ago | parent | next [-] |
| I like this so much. Maybe I’m romanticizing but hoping tools like this give local news a chance against monopolies like Sinclair Broadcasting. |
| |
| ▲ | cyanydeez 2 days ago | parent [-] | | It'll merely generate hyper local bots that generate hyper local slop. The media landscape is downhill. | | |
| ▲ | ChrisMarshallNY 3 hours ago | parent | next [-] | | I think it's just lead generation. If the paper decides to use AI to write/edit stories, that's another topic. | |
| ▲ | jasonlotito 2 hours ago | parent | prev | next [-] | | That's a long way of saying you didn't read the article. | |
| ▲ | binarymax 7 hours ago | parent | prev [-] | | Are search summaries slop? | | |
| ▲ | phoghed 3 hours ago | parent | next [-] | | Too many people use "slop" as a synonym for "AI Generated", so it's becoming useless as a term | | | |
| ▲ | ginko 7 hours ago | parent | prev [-] | | Yeah | | |
| ▲ | binarymax 6 hours ago | parent [-] | | It’s not a blanket ‘Yeah’. It’s a ‘sometimes’ and that sometimes depends on use. Distilling down a whole days worth of research is one of the best uses of AI. | | |
| ▲ | irishcoffee 4 hours ago | parent [-] | | My 12 year old told me that their friend group says "Oh I AI'd that" when they make a mistake, as a joke. Yet you trust AI to distill down a days worth of information in a 100% accurate way? When pre-teens already understand how unreliably they are? | | |
| ▲ | binarymax 2 hours ago | parent | next [-] | | If I'm getting citations, and the distillation has strong alignment with citations, then it's fine. In other words: the if the summary is a highlight and not a rewrite then I'm happy. I've studied this extensively and have even created metrics around this problem: https://maxirwin.com/articles/llm-rag/ | | |
| ▲ | irishcoffee an hour ago | parent [-] | | So, if you review the entire dataset? At that point just write the summary yourself. I use "AI" a lot, it is a fantastic tool. We need to stop pretending it is some kind of panacea. It's a tool. | | |
| ▲ | Forgeties79 an hour ago | parent [-] | | > I use "AI" a lot, it is a fantastic tool. We need to stop pretending it is some kind of panacea. It's a tool. I have been repeating some variation of this for months. If people would stop acting like it’s THE tech solution to ALL things ALL the time I bet a lot of critics would quiet down. The overhyping has become exhausting. It’s been going on for years. GPT messed up the math for me the other day when I was simply adding 10 durations for a TRT. Couple of HH:MM:SS inputs, annoying to add up and I had it open. It got it wrong, I told it it was wrong, it got it wrong again. Then out of curiosity I provided it the answer, asked for it to confirm it against the original numbers, and it went “you’re absolute right, it’s [original/wrong answer from earlier].” This stuff happens probably 10-15% of the time for me regardless of the model. Not just math, just super simple crap. It’s wild to see at this point after 3-4 solid years of “hyperscaling” and overhyping. And it’s the kind of thing that keeps people like me from buying in beyond the foot or two we’ve stuck in the water. |
|
| |
| ▲ | Zambyte 3 hours ago | parent | prev | next [-] | | What does "100% accurate" distillation even mean? That sounds contradictory. "Good enough" is fundamentally what makes distillation valuable, not perfection. | |
| ▲ | ToucanLoucan 3 hours ago | parent | prev [-] | | I will say, while I agree with the broad strokes of what you're saying, in my experience when you provide an LLM with data to summarize, instead of it needing to go find it, the hallucination rate goes down staggeringly. I've never had a huge hallucination when the material I'm asking to be summarized is provided to it at the time of the request. Any further distance than that though, even summarizing the conversation it's aware of up to that point, is dicier. |
|
|
|
|
|
|
|
| ▲ | ChrisMarshallNY 3 hours ago | parent | prev | next [-] |
| Pretty cool. I guess the downside might be, that someone could be replaced by this, but it seems that this is a perfectly valid application of AI. In fact, lead generation seems to be a natural fit for LLMs. |
|
| ▲ | mmooss 3 days ago | parent | prev | next [-] |
| > Early on, the team tried to balance quality with the high cost of searching many small, scattered sources. They discovered that prompting can implicitly control search depth and behavior, but only after trial and error. I don't understand the cost here. The number of sources should be relatively tiny - it's not a general Internet search. The amount of data scraped would seem to be relatively tiny - how much data is there on suburban-county, PA? |
| |
| ▲ | mrweasel 7 hours ago | parent | next [-] | | For the individual county it's probably not a ton of data, but multiply it up, it's a lot for all the counties in e.g. Philadelphia. Normally there is data, meetings, events and so that you wouldn't report on in a newspaper, because it's really only relevant to maybe a few thousand people, but for those thousands it might be really important. My city publishes a lot of stuff on their website, it's hard to find, relevant to maybe a thousand people, maybe less. The school board meetings, public utility companies, companies in general all publish massive amounts of information that's never surfaced, but is relevant to those living in the vicinity. Being able to collect all of this, sort it, assess it's relevans and produce hyper local news could be a massive boost for local grassroots movements and participation in local affairs and elections. | | |
| ▲ | thephyber 7 hours ago | parent [-] | | Agree that the permutation becomes a problem, but the organization should be super simple: the state mandates the use of some standard index file format for all municipalities governments, not unlike how sitemap.xml works for a website index. There is a community-run project that standardizes each election's results across states, counties, and precincts. It could be a template for local info releases. | | |
| ▲ | nemomarx 3 hours ago | parent [-] | | do you have documentation on that standard index file? I'm not sure I've seen a Pennsylvania gov page explaining it |
|
| |
| ▲ | tclancy 7 hours ago | parent | prev [-] | | I think it’s more about how many different types of content you’re looking at. Public google calendars, Facebook event pages, personal sites, etc. |
|
|
| ▲ | FL410 5 hours ago | parent | prev | next [-] |
| >The result: Scrape evolved from a small experiment to “load-bearing and critical infrastructure” Please be satire |
| |
|
| ▲ | dec0dedab0de 4 hours ago | parent | prev | next [-] |
| I'm still annoyed at them for shutting down philly.com |
|
| ▲ | calvinmorrison 7 hours ago | parent | prev [-] |
| http://www.mincedgarlic.org/inky.png |