CalvusRex

u/CalvusRex@lemmy.world
3 posts · 8 comments

Recent posts

Recent comments

on Future RAG system fodder · c/history · 4 pts · 22d

@queerlilhayseed@piefed.blahaj.zone Thanks for the positive response! You basically named the two hurdles that ate most of my development time on my projects system.

On the counting/nondeterminism thing: the trick was giving up on the LLM as a "knower" entirely and treating it as nothing more than a reader. So if I want to know how many times grass shows up in some sprawling Tolkien-esque passage, I'm not asking the model "how many times does grass appear" and hoping for the best. I use it to pull out and tag the relevant sentences into a structured database, and then a dumb, boring, deterministic Python script does the actual counting. Temperature's pinned at 0, and I force it to hand back verbatim quotes with citations instead of letting it synthesize a summary. That alone killed most of the hallucination problem.

The "systemic wrongness" side (bias, basically the model just inventing stuff) came down to hybrid search, meaning keyword/BM25 alongside vector embeddings. Pure semantic search will absolutely whiff on a weirdly named character or a specific term just because it doesn't "feel" semantically close to the query, so having the keyword layer as a backstop matters a lot. And I'm strict about grounding: no direct source quote from retrieval means the system says "I don't know" instead of guessing. It's not allowed to fill gaps with vibes.

Your instinct about cataloging prompts was right on too. I keep a golden dataset of queries I run every time I touch the system, and rather than some vague overall correctness score, I bucket failures by type: was it a retrieval failure or a synthesis failure. That distinction alone tells you whether the fix is in your chunking strategy or your prompting, instead of just flailing at both.

Honestly the tools have come a long way since you last looked at this stuff, and I think it's mostly a philosophy shift. People stopped treating LLMs like databases and started treating them like interfaces sitting in front of one.

so i’ve seen some great answered here. Some folks are leery of some companies but so far i have found that Tailscale and Cloudflare a great combination. My ISP doesn’t let you port forward and TS & CF circumvent that nicely. Zero issues, both free services and saves me a ton of headache.