Building semantic search: batches, shifting scores and cheaper retrieval
What a small Signals experiment taught us about Jev relevance scores, batch context and lower-cost candidate retrieval.
By Clairevue · · 5 min read

TL;DR
- We use Jev to judge whether posts answer a search question, rather than relying on matching keywords.
- The same post received different scores when its neighboring posts changed. Identical requests varied too, though less in this test.
- Vector retrieval followed by Jev cut query-time usage by about 84%, but recovered only ten of 21 original qualifying matches.
- Our default cutoff is now 85%, down from 90%. Hybrid retrieval and broader, explicitly paid searches are possible next steps—not established solutions.
How we used Jev
Clairevue Signals collects AI-related posts: product releases, jobs, events and other developments. A useful search should understand requests such as “AI funding announcements rather than hiring posts,” even when the wording differs from the posts.
We built this with Jev. Each post gets a separate true/false relevance question, using the API's noul response type. We use its returned value, between zero and one, as the relevance score. The instructions explicitly tell Jev to judge only the specified post. Our recorded responses contained no separate confidence field.
Sending hundreds of full posts in one request hit the provider's context limit. The working design sends up to 64 posts per request, using the first 1,000 characters of each post plus selected metadata. Two requests run in parallel; further batches follow as a slot becomes free, up to eight requests per search.
That covers at most 512 recent eligible posts, after the visitor's filters. It keeps requests manageable, but neither uncapped source text nor the whole historical collection is searched. Applying tags first narrows the work and makes that window more useful.
The same post, different scores
A cheaper-search experiment exposed a problem with our original 90% display cutoff. Three posts that qualified when Jev judged the full candidate set fell below the cutoff when it judged a smaller shortlist. Their scores moved from 92% to 89%, and from 90% to 86% for two vacancies.
We then held one post fixed: a Textlayer AI Architect vacancy. The question asked for AI engineering jobs explicitly located in Canada. The post said applicants must be eligible to work in Canada—a geographic clue, though not an unambiguous job location.
We placed it in a ten-post batch with nine other relevant jobs, then replaced one neighbor at a time with an irrelevant job. The target's text, metadata, position and question stayed unchanged. Replacement lengths were closely matched. Each of the ten batches was sent twice, in shuffled order.
With the original nine neighbors, the target scored 82% and 84%. With all nine replaced, it scored 94% both times. Its average rose by eleven percentage points, although the progression was not strictly upward at every step.
Identical requests also differed: six of ten duplicate pairs changed, by up to three percentage points. Two pairs straddled our cutoff: 89% versus 90%, and 90% versus 87%.
This supports batch-context sensitivity alongside request-level variation in this one case. It does not establish how Jev works internally or how common the effect is across tasks. Two repeats cannot provide a reliable estimate of general instability; hidden provider caching also remained unresolved.
We have lowered the default to 85%, allowing more borderline posts through. On the saved shortlist scores, that would retain the three posts lost at 90%. It may also admit weaker results. It does not cure variability: the fixed target still fell below 85% in the initial ten-post batch. The new cutoff is a product choice, not a freshly validated accuracy threshold.
Trying cheaper retrieval with Polygres
We also tested a different way to select candidates. Polygres supplied the retrieval infrastructure; OpenAI's text-embedding-3-small supplied numerical representations of the posts, called embeddings. Vector search finds representations close to the question, reusing that preparation instead of asking Jev to read every candidate again.
The comparison used a frozen copy of 527 posts and twelve questions, separate from the production window. An AI assistant assessed relevance before search outputs; a human reviewed 56 assessments. These judgments were our scorecard, not independently verified accuracy or a filter applied to the answers.
We ran actual Polygres top-50 retrieval followed by Jev judgment, preserving the original filters. Jev scored 505 question–post pairs rather than 3,429. Across twelve questions, query-time usage fell from about 5.12¢ to 0.80¢—roughly 84% less. Including the initial corpus embeddings brought the combined approach to about 1¢. These are documented-rate estimates and included-allowance consumption, excluding labor and hosting costs.
The savings came with missed answers. At the original 90% cutoff, only thirteen of Jev's 21 qualifying matches entered the shortlist. Three more fell below the cutoff when rescored, leaving ten of the original 21. Seven funding announcements were among the posts retrieval excluded. Lowering a cutoff cannot recover posts Jev never receives.
Across the eight original comparison questions reserved from development, the assessed-useful share of the first ten results was 35% for Jev alone and 28.75% for the combined approach. The follow-up reused already-seen questions, so it was exploratory rather than fresh validation.
Indexed retrieval offers a route to searching more than a recent window at low query cost. This particular setup did not establish equivalent relevant coverage. Polygres is not a competing decision model, and this experiment cannot settle vector search generally.
What we would try next
First, combine vector and lexical retrieval: meaning-based candidates plus candidates found through words, names and phrases. Merge them before judging relevance. Our earlier planned hybrid request was rejected by the installed tool's argument validation, so we have no measured hybrid result yet. A larger shortlist, new questions and stability tests on more targets would help assess the next attempt.
Second, make broader search an explicit choice. Paging through an already-scored result list is free of new model calls; searching another candidate window is different. It adds coverage and paid evaluation, with possible score changes as batches change. Clear “search older posts” controls could expose that trade-off rather than imply every query searches everything.
Paid accounts could fund a fuller-corpus search with visible scope, usage limits and costs. More budget would allow more requests, not remove the per-request context limit or guarantee stable scores. These are proposed product options, not implemented features.
For now, we have a working bounded search and evidence of where its decisions become fragile. The next task is to preserve useful candidates cheaply while making the search window and the uncertainty around borderline scores clear.