| ▲ | Terr_ 7 hours ago | ||||||||||||||||
> this is NOT something most people want to do I've always felt that "proper" and responsible LLM use for search [0] would be two boxes: You can describe what you want in the first, and it'll propose search terms in the second, and then those get executed normally. Yes, the "average user" [1] might not usually care about the second box... Until they need to because the query/results are wrong. Showing them in tandem means: 1. Users are at least capable of learning through exposure. 2. Users may realize a key term can be added which the model could never have guessed. 3. Users may recognize a term in there that doesn't make sense, allowing them to detect a translation error. 4. If good search-terms leads to a bad outcome, it is possible for someone to report and diagnose it, rather than a fully black-box mystery. _____ [0] Not just for websites, but also things like internal business software, or SQL queries. [1] The average that might not exist. ( https://www.thestar.com/news/insight/when-u-s-air-force-disc... ) There are some features that everybody needs, just at different times. | |||||||||||||||||
| ▲ | savory_pancake 4 hours ago | parent | next [-] | ||||||||||||||||
In the words of Edsger Dijkstra, "Projects promoting programming in 'natural language' are intrinsically doomed to fail." I think for something as bespoke as a search engine, a DSL that accepts quotes, logical operators, and specific terms definitely allows for far greater specificity than natural language can (at least in an equivalent amount of text). I think the two-tiered input-output approach you propose makes sense. Allowing users to inspect and mutate lower-level languages allows the user to make modifications as needed. I think it's very much in the spirit of free software. But then again, for text-based search specifically (search engines, notes, etc.), I think there is some value in querying the LLM directly, as it is able to fuzzy search by inspecting its weights. This results in a lesser degree of specificity, which allows for more false positives, but can maybe capture similar words (i.e., synonyms / typos / tenses) or higher-level semantic concepts. Maybe they're just two different search algorithms, and the user should be able to choose between them. Thanks for linking the article, I found it very interesting. | |||||||||||||||||
| ▲ | tokioyoyo 6 hours ago | parent | prev [-] | ||||||||||||||||
Your described thinking pattern does not apply to 95%+ users, and might add confusion leading to retention loss (e.g. 2 text boxes? Wtf, two boxes? What do i type in the second one? Whatever i’ll switch to the usual). Every project with at least 100K non-tech users I’ve worked on, has shown that sort of thinking just doesn’t translate. And regarding “average doesn’t exist” - that’s true. But no company does a/b testing to land on average. I’d assume a good 85%-kinda pass rate for these type of experiments. | |||||||||||||||||
| |||||||||||||||||