Orchestrate

Hybrid search and metadata filtering: keywords, vectors and the filters that keep answers in scope

Why keyword and vector search fail in different places, how to fuse them with RRF or weighted scores, and how metadata filters add precision and enforce access.

Updated Oct 4, 2026Foundational 8 minDeep 40 min
Foundational layer · 8 min read

The 60-second version

An AI assistant that answers from company documents has to find the right passages first. There are two classic ways to search.

Keyword search looks for the words in the question. It is excellent at exact things: an error code like PRC-112, a purchase order number, a material number. It is poor at paraphrase: "my customer's order is stuck" won't find a note that says "sales order blocked by the credit check".

Vector search compares meaning. It finds the "blocked" note for the "stuck" question. But it tends to blur rare codes and numbers, which carry little meaning on their own.

Hybrid search runs both and merges the two ranked lists. Each one covers the other's blind spot.

Metadata filters come first. Before ranking anything, the search keeps only passages with the right labels: the company codes this user may see, the right language, documents valid today. Filters make answers more precise, and, when set from the user's real permissions, they keep answers inside what the user is allowed to know.

Why it matters to the business

Retrieval quality decides answer quality. A wrong or missing passage produces a confident wrong answer, and nobody sees the passage that wasn't found.

Take the running order-to-cash example. A credit team asks the assistant about blocked sales orders.

  • The wrong policy. The company changed its credit release rule in July. Both the old and the new policy mention "credit", "limit" and "release". Without a "valid on today's date" filter, the old policy can rank first, and the assistant quotes a rule that no longer applies.
  • The wrong company. A user in company code 1000 asks about a customer. A case note from company code 3000, which the user may not see in SAP, is the closest match. Without an access filter, it lands in the answer. That is a data leak, not a quality issue.
  • The missed code. A user pastes the error code from the screen. Pure vector search returns notes about pricing in general and misses the one note that explains that exact code.

The same pattern shows up in procure-to-pay (three-way match exceptions with purchase order numbers) and plan-to-produce (MRP exceptions with material numbers). SAP work is full of codes, numbers and organizational units. That is why hybrid search and filters matter more here than in a generic chatbot.

The cost is modest. Keyword search is cheap and mature. Filters can make a search faster or slower depending on the engine, so they need measuring, as Vector databases explained showed. The main investment is labelling every document with the right metadata, and testing.

How SAP does it

As of October 2026, SAP's two retrieval options handle this differently.

  • SAP AI Core grounding (generative AI hub). SAP's release notes of 16 February 2026 added advanced filtering, including metadata filtering, to vector search. The same release lets you manage metadata on documents, collections and chunks, and merge and rank results from several data repositories with the Retrieval API. SAP's SDK shows a grounding request that filters documents by a metadata key and value. This course found no SAP statement that the grounding service offers keyword or hybrid ranking; check with SAP for your release.
  • SAP HANA Cloud. Vectors sit in ordinary tables, so a similarity search can carry an ordinary SQL WHERE condition on company code, language or dates. SAP's open-source langchain-hana package adds filter operators, including a keyword filter. To fuse keyword and vector rankings, you can run both queries and merge the lists in your application, as the deep layer shows.

Either way, the filter values must come from the user's SAP authorizations, applied on the server. A later topic in Unit 7 covers grounding on SAP data without breaking authorizations.

Which search for which question

Question type Example Keyword Vector Hybrid
Exact code or number "PRC-112", "PO 4500017311" Strong Weak Strong
Paraphrase or synonym "order stuck" vs "order blocked" Weak Strong Strong
Specialist jargon "three-way match tolerance" Strong Good Strong
Other language German note, English question None Possible with a multilingual model Possible
Must stay in scope Only my company codes, only valid policies Needs filters Needs filters Needs filters

Two rules of thumb follow:

  1. Start hybrid by default for SAP content, because users mix codes and plain language in one question.
  2. Treat filters as two separate things. Precision filters (language, document type, valid date) improve answers and may be relaxed. Access filters (company code, sales organization, confidentiality) protect data and must never be relaxed.

Questions to ask

  • Which questions do users actually ask? How many contain codes, numbers or names?
  • Have we measured keyword-only, vector-only and hybrid on our own test questions?
  • How are the two result lists merged, and who tuned the settings?
  • Which metadata does every chunk carry: company code, sales organization, plant, document type, language, valid-from and valid-to dates?
  • Where do the access filter values come from? Are they derived from the user's SAP roles on the server, or sent by the browser?
  • What happens to a document that is missing a label? Is it excluded or shown to everyone?
  • How do superseded documents leave the results?
  • If we use SAP AI Core grounding, which metadata filters does our release support, and what happens when a key is missing?

Common misconceptions

  • "Vector search replaced keyword search." Embeddings handle meaning well and exact codes poorly. Most production search stacks combine both.
  • "Hybrid search is always better." Usually, not always. Fusion can push a strong single result down. Measure it on your questions.
  • "A filter in the prompt is enough." Telling the model to "ignore other company codes" is not a control. Leave the passages out of retrieval altogether.
  • "Filters are free." Depending on the engine, a filter can speed a search up or slow it down, and very selective filters are the hard case. Vector databases explained measured this.
  • "Missing labels are harmless." A document without a company code is either invisible or visible to everyone. Decide which, and prefer invisible.

Key terms

  • Keyword (lexical) search: ranking by the words a document shares with the question.
  • BM25: the standard keyword ranking formula; it rewards rare words and repeated matches, with limits.
  • Vector (semantic) search: ranking by closeness of embeddings, which represent meaning.
  • Hybrid search: running keyword and vector search and merging the results.
  • Fusion: the rule that merges two ranked lists into one.
  • RRF (reciprocal rank fusion): a fusion rule that uses only each document's rank in each list.
  • Metadata: labels stored with each passage, such as company code, language or valid-from date.
  • Precision filter: a filter that narrows results to what is relevant.
  • Access filter: a filter that limits results to what the user is allowed to see.
  • Fail closed: when information is missing, deny rather than allow.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1A user pastes error code PRC-112 into the assistant and gets general pricing advice instead of the note about that code. What is the likely cause?

    Answer: B. Exact codes carry little meaning for an embedding, so vector search tends to miss them. Keyword search matches the exact string, which is why hybrid search covers both cases.
  2. 2What does hybrid search add over keyword search or vector search alone?

    Answer: C. Keyword search is strong on codes and names, vector search on paraphrase. Merging the lists covers both, but it can still push a strong single result down, so measure it.
  3. 3The assistant quotes a credit release rule that changed in July. Which fix fits best?

    Answer: D. The old and new policies use the same words, so ranking alone can't separate them. A valid-from and valid-to filter removes the superseded policy before ranking.
  4. 4Which statement about access filters is right?

    Answer: A. An access filter protects data, so it is set from who the user is and never relaxed. Precision filters, such as language, may be relaxed; access filters may not.
  5. 5A chunk has no company code label. What should the search do?

    Answer: B. Failing closed means a missing label denies access rather than granting it. Otherwise any unlabelled document becomes visible to every user.
  6. 6A partner proposes vector-only retrieval for SAP procure-to-pay notes. What should you ask first?

    Answer: D. Procure-to-pay questions are full of purchase order and material numbers, where vector search is weakest. Testing those questions shows whether keyword or hybrid search is needed.
  7. 7Your team uses SAP AI Core grounding. What has SAP documented for filtering as of 2026?

    Answer: C. SAP's release notes of 16 February 2026 added metadata filtering to vector search and metadata management in the Vector API. The course found no SAP statement on hybrid ranking or automatic role-based filters.
Deep layer · 40 min read

Mental model

Retrieval is two jobs: decide what may be searched, then decide what ranks first. Filters do the first job. Rankers and fusion do the second. Never let the second job do the first.

  • Filters are set membership. A passage passes or it doesn't. Access filters (company code, sales organization, confidentiality) come from the user's identity on the server. Precision filters (document type, language, validity date) come from the question or the use case.
  • Rankers have different blind spots. BM25 matches strings and can't see synonyms. Embeddings see meaning and blur rare strings such as PRC-112 or 4500017311.
  • Fusion merges ranked lists whose scores don't share a scale. A BM25 score of 3.8 and a cosine similarity of 0.53 can't be added directly. You either fuse by rank (RRF) or rescale scores first (weighted fusion).

In RAG fundamentals retrieval was one box. In Vector databases explained you opened the vector index. This topic adds the keyword ranker, the fusion step and the filter step around it.

How it works

flowchart LR
  U[User identity] --> F[Access filter<br/>company codes]
  Q[Question] --> P[Precision filters<br/>type, language, date]
  F --> S[Allowed passages]
  P --> S
  S --> K[BM25 ranking]
  S --> V[Vector ranking]
  K --> R[Fusion<br/>RRF or weighted]
  V --> R
  R --> T[Top passages to the LLM]

Keyword search with BM25

BM25 scores a document for a question by adding up a weight for every question word the document contains. Three ideas sit inside the formula:

  • Rare words count more. The inverse document frequency (IDF) of a word is high when few documents contain it. prc-112 appears once in the lab's notes; order appears in many. A match on prc-112 dominates.
  • Repeats help, with diminishing returns. The k1 parameter controls how fast extra occurrences of a word stop adding score. Elasticsearch's default is k1 = 1.2.
  • Long documents are normalised. A long document contains more words by chance. The b parameter (default 0.75 in Elasticsearch) controls how much the score is scaled down for long documents.

For one question word, the lab's code computes idf × tf × (k1 + 1) / (tf + k1 × (1 − b + b × length / average_length)), where tf is how often the word appears in the document. BM25 is Elasticsearch's default similarity, and Microsoft's Azure AI Search uses it for the keyword half of hybrid queries.

Tokenization matters as much as the formula. The lab keeps hyphenated codes such as rm-4711 as one token. Split them into rm and 4711, and a match on 4711 alone may hit unrelated documents. Keyword engines solve this with analyzers; whichever you use, test your codes.

The question and each passage become embeddings, and the score is cosine similarity, as in Embeddings and semantic similarity. Two properties matter for fusion:

  • Vector search always returns something. Every passage has a similarity, even when nothing is relevant. Keyword search can return nothing, which is a useful signal.
  • Scores cluster. Cosine similarities for one question often sit in a narrow band, so a difference of 0.05 may separate the best match from noise.

Fusion 1: reciprocal rank fusion (RRF)

RRF ignores the scores and uses only positions. Each list adds 1 / (k + rank) for every document it contains, with rank counting from 1. Documents that both lists rank highly win.

Elasticsearch documents the formula with rank_constant (k) defaulting to 60, and states that RRF "requires no tuning", because the two scoring systems never have to be compared. The method comes from Cormack, Clarke and Buettcher (SIGIR 2009), which Elasticsearch cites.

The constant changes behaviour. With k = 60, rank 1 earns 1/61 and rank 2 earns 1/62: nearly equal. Agreement between the lists matters much more than position within one list. With a small k, the top of each list matters more. Defaults differ between products: Qdrant documents k = 2 as its default (configurable since version 1.16.0). Don't assume two products fuse the same way.

Fusion 2: weighted score fusion

Rescale each list's scores to the range 0 to 1, then combine: alpha × vector + (1 − alpha) × keyword. Weaviate uses this design: alpha runs from 0 (keyword only) to 1 (vector only), with a default of 0.75, and its relativeScoreFusion method (the default since Weaviate 1.24) scales scores relative to each result set before combining. Qdrant documents another variant, distribution-based score fusion (DBSF), which normalises using the mean and three standard deviations of each list's scores.

Weighted fusion keeps score gaps that RRF throws away: a vector match far ahead of the rest stays far ahead. The price is a parameter to tune and edge cases to handle. In the lab, a list where every score ties (for example, a vector search where the question has no known meaning) must not turn into "every document scores 1.0".

RRF Weighted score fusion
Uses Ranks only Rescaled scores
Tuning k (and the window size) alpha and the rescaling method
Strength Robust, no score calibration Keeps "how much better" information
Weakness A strong single-list hit ranks behind documents both lists like Sensitive to outliers and to ties
Seen in Elasticsearch, Azure AI Search, Qdrant Weaviate, Qdrant (DBSF)

Metadata filters

A filter is a condition on the labels stored with each chunk. Typical SAP labels:

Label Example values Kind of filter
Company code 1000, 2000, or ALL for shared content Access
Sales organization, plant 1010, 1710 Access or precision
Confidentiality internal, confidential Access
Document type policy, case_note, help Precision
Language en, de Precision
Valid from / valid to 2026-07-01 / 9999-12-31 Precision

Most stores express filters as a small query language. SAP's langchain-hana package, for example, accepts operators such as $eq, $in, $between, $like and $contains, combined with $and and $or. A filter for "company code 1000 or shared, valid on 1 October 2026" looks like this in that style:

# Shape of a metadata filter in the operator style langchain-hana documents (not run here).
# Dates are stored as YYYYMMDD integers so the range comparison works.
where = {"$and": [
    {"company_code": {"$in": ["1000", "ALL"]}},
    {"valid_from": {"$lte": 20261001}},
    {"valid_to": {"$gte": 20261001}},
]}

Two rules for hybrid search:

  1. Apply the same filters to both rankers. If the keyword half is unfiltered, fusion can pull a forbidden document back into the results. Microsoft documents that in Azure AI Search a filter applies to both halves of a hybrid query; check your own stack.
  2. Filter before you rank, wherever the engine allows it. Filtering after ranking can return fewer results than asked, and a forbidden passage is still read by the ranker. How engines combine filters with an approximate index is covered in Vector databases explained.

Filters as an authorization tool

OWASP lists "vector and embedding weaknesses" among the top risks for LLM applications in 2025. It warns that misaligned access controls let embedded data leak, including across groups that share one store, and recommends fine-grained access controls and permission-aware vector stores. A metadata filter is the usual mechanism. It only works if:

  • The values come from the server. Look up the user's company codes from their identity and roles. Never accept a filter from the browser or from the language model.
  • Labels are complete. A missing label must exclude the chunk (fail closed).
  • Labels are right at ingestion. The chunk inherits labels from its source document. A wrong label is a wrong permission.
  • Retrieval is logged. Record who searched, the filter applied and which chunks came back, so you can audit a leak.

Build it yourself: a hybrid search lab with filters

Before you start: complete Set up your computer for this course and Set up for Unit 7. They create the orchestrate-course folder, its .venv and the unit07 folder. For real embeddings you also need the Unit 3 model from Set up for Unit 3; without it, use --offline.

You will build hybrid_search.py. It holds 16 made-up SAP help notes and case notes, each labelled with a company code, document type, language and validity dates. It ranks them with BM25, with vector search, and with two kinds of fusion, after applying an access filter for the signed-in user and optional precision filters. An eval command scores every method on 10 test questions, so you can see each method's blind spot in numbers.

flowchart LR
  N[16 labelled notes] --> F{Filters<br/>user + options}
  F --> B[BM25]
  F --> E[Embeddings]
  B --> R[RRF / weighted]
  E --> R
  R --> O[Top 3 notes]
  T[10 test questions] --> V[eval: hit@3 and MRR<br/>per method]

What you need

  • Your course folder with the Unit 7 setup done. Cost: free.
  • No new library. The script uses numpy (installed with Chroma in Unit 7) and Python's built-in modules. The keyword ranker is written out in full, so you can read every line.
  • Optional: sentence-transformers and the all-MiniLM-L6-v2 model from Unit 3 for real embeddings. Without them, add --offline.
  • About 45 minutes. No account and no network with --offline.

Step 1: Open your course folder

  1. Open a terminal (on Windows, PowerShell; on macOS, Terminal) and turn on your environment.

    Windows (PowerShell):

    cd $HOME\orchestrate-course
    .\.venv\Scripts\Activate.ps1

    macOS/Linux:

    cd ~/orchestrate-course
    source .venv/bin/activate
  2. Check that numpy is there:

    pip show numpy

    You should see Name: numpy and a Version: line. If you see WARNING: Package(s) not found, add numpy to requirements.txt and run pip install -r requirements.txt.

Step 2: Create the script

  1. In VS Code, right-click the unit07 folder, choose New File, name it hybrid_search.py, paste the code below and save.
"""Unit 7: hybrid search with metadata filters over SAP-style help notes.

It ranks the same notes three ways: keyword search (BM25), vector search (embeddings), and a
fusion of the two (reciprocal rank fusion or a weighted score). Before any ranking, it applies
metadata filters: the company codes the user may see (an access rule, set on the server side)
and optional precision filters (document type, language, valid on a date).

How to run (from your course folder, with .venv turned on):
    python unit07/hybrid_search.py search "why is the customer's order stuck"
    python unit07/hybrid_search.py search "fix for PRC-112" --mode all
    python unit07/hybrid_search.py search "credit limit release" --user anna --as-of 2026-10-01
    python unit07/hybrid_search.py eval                    # hit@3 and MRR for every method
Add --offline to any command to use toy "concept" embeddings instead of the Unit 3 model.
"""
import argparse
import math
import re
import sys
from collections import Counter
from datetime import date

try:
    import numpy as np
except ImportError:
    sys.exit("numpy is not installed. Run: pip install -r requirements.txt")

MODEL = "sentence-transformers/all-MiniLM-L6-v2"  # the Unit 3 model; 384 numbers per text

# ---------------------------------------------------------------------------
# Made-up help notes. Metadata sits next to the text, as it would in a vector store.
# company_code "ALL" means the note applies to every company code.
# A note WITHOUT company_code is a data error: the filter must leave it out (fail closed).
# ---------------------------------------------------------------------------
NOTES = [
    {"id": "N01", "company_code": "1000", "doc_type": "policy", "language": "en",
     "valid_from": "2025-01-01", "valid_to": "2026-06-30",
     "text": "Credit policy v1: a sales order blocked by the credit check may be released by the "
             "credit analyst if the exposure exceeds the limit by less than 5 percent."},
    {"id": "N02", "company_code": "1000", "doc_type": "policy", "language": "en",
     "valid_from": "2026-07-01", "valid_to": "9999-12-31",
     "text": "Credit policy v2: a sales order blocked by the credit check may be released by the "
             "credit analyst only if the exposure exceeds the limit by less than 2 percent; above "
             "that the credit manager approves."},
    {"id": "N03", "company_code": "1000", "doc_type": "case_note", "language": "en",
     "valid_from": "2026-03-12", "valid_to": "9999-12-31",
     "text": "Customer order stuck for three days. Cause: credit exposure above the limit after an "
             "unpaid invoice. Released after the payment arrived."},
    {"id": "N04", "company_code": "2000", "doc_type": "case_note", "language": "en",
     "valid_from": "2026-05-02", "valid_to": "9999-12-31",
     "text": "Customer order held because the credit limit was exceeded. The analyst raised the "
             "limit after the sales manager approved."},
    {"id": "N05", "company_code": "ALL", "doc_type": "help", "language": "en",
     "valid_from": "2024-01-01", "valid_to": "9999-12-31",
     "text": "Error PRC-112 in pricing means a mandatory price condition is missing. Maintain the "
             "condition record for the customer and material, then redetermine prices."},
    {"id": "N06", "company_code": "ALL", "doc_type": "help", "language": "en",
     "valid_from": "2024-01-01", "valid_to": "9999-12-31",
     "text": "Pricing errors on the order: check that the price, discount and tax conditions are "
             "maintained and valid on the pricing date."},
    {"id": "N07", "company_code": "1000", "doc_type": "case_note", "language": "en",
     "valid_from": "2026-08-20", "valid_to": "9999-12-31",
     "text": "Invoice quantity differs from the goods receipt for purchase order 4500017311. "
             "Supplier delivered 90 of 100 pieces; invoice blocked for payment."},
    {"id": "N08", "company_code": "ALL", "doc_type": "help", "language": "en",
     "valid_from": "2024-01-01", "valid_to": "9999-12-31",
     "text": "Three-way match: the invoice is checked against the purchase order and the goods "
             "receipt. A quantity or price variance above tolerance blocks the invoice for payment."},
    {"id": "N09", "company_code": "2000", "doc_type": "case_note", "language": "en",
     "valid_from": "2026-04-11", "valid_to": "9999-12-31",
     "text": "Vendor billed a higher price than agreed on the purchase order. Price variance above "
             "tolerance; buyer contacted the vendor for a credit memo."},
    {"id": "N10", "company_code": "1000", "doc_type": "case_note", "language": "en",
     "valid_from": "2026-09-03", "valid_to": "9999-12-31",
     "text": "Planning run showed a shortage of component RM-4711 for production in week 38. "
             "Planner moved the order forward and expedited the supplier."},
    {"id": "N11", "company_code": "ALL", "doc_type": "help", "language": "en",
     "valid_from": "2024-01-01", "valid_to": "9999-12-31",
     "text": "MRP exceptions: the planning run flags missing parts, late receipts and orders to "
             "reschedule. Review exceptions daily, starting with the critical components."},
    {"id": "N12", "company_code": "2000", "doc_type": "case_note", "language": "en",
     "valid_from": "2026-07-15", "valid_to": "9999-12-31",
     "text": "Missing parts for assembly line 2: component RM-5120 not delivered. Planner changed "
             "the procurement to a second supplier."},
    {"id": "N13", "company_code": "3000", "doc_type": "case_note", "language": "en",
     "valid_from": "2026-06-01", "valid_to": "9999-12-31",
     "text": "Subsidiary customer order blocked: credit limit exceeded after a merger. Limit "
             "reviewed by the regional credit manager."},
    {"id": "N14", "company_code": "1000", "doc_type": "help", "language": "de",
     "valid_from": "2024-01-01", "valid_to": "9999-12-31",
     "text": "Kundenauftrag gesperrt: das Kreditlimit ist überschritten. Der Kreditanalyst prüft "
             "das Obligo und gibt den Auftrag frei."},
    {"id": "N15", "company_code": "ALL", "doc_type": "help", "language": "en",
     "valid_from": "2024-01-01", "valid_to": "9999-12-31",
     "text": "Payment terms: the due date of an invoice is calculated from the baseline date and "
             "the payment terms of the customer or supplier."},
    {"id": "N16", "doc_type": "case_note", "language": "en",
     "valid_from": "2026-09-30", "valid_to": "9999-12-31",
     "text": "Confidential: customer order blocked; credit limit cut after a rating downgrade. "
             "Board approval needed for release."},
]

# Who may see which company codes. In a real system this comes from the user's SAP roles,
# looked up on the server from the signed-in identity, never from the request.
USERS = {
    "anna": ["1000"],
    "ben": ["1000", "2000"],
    "chen": ["1000", "2000", "3000"],
}

# Test questions with the note a good search should rank first. Used by `eval`.
EVAL_SET = [
    # exact identifiers: keyword search should shine
    ("PRC-112", "ben", "N05"),
    ("4500017311", "ben", "N07"),
    ("RM-4711 status", "ben", "N10"),
    ("RM-5120", "ben", "N12"),
    # paraphrases with few shared words: vector search should shine
    ("client purchase frozen because they had not paid", "anna", "N03"),
    ("supplier sent fewer pieces than we were billed for", "ben", "N07"),
    ("vendor charged more than agreed", "ben", "N09"),
    ("lacking components to build product", "ben", "N12"),
    # a mix of both
    ("who may release a credit hold above 3 percent", "anna", "N02"),
    ("what does error PRC-112 mean", "ben", "N05"),
]

STOPWORDS = set("a an the is are was of for to in on by and or if what why who how when does "
                "do may be it its that this with after from than our my me i s".split())


def tokens(text):
    """Lowercase words; keeps codes such as prc-112 and 4500017311 as one token."""
    return [t for t in re.findall(r"\w+(?:-\w+)*", text.lower()) if t not in STOPWORDS]


# ---------------------------------------------------------------------------
# Keyword search: BM25
# ---------------------------------------------------------------------------
class BM25:
    """Okapi BM25 with the common defaults k1=1.2 (term-frequency saturation) and b=0.75
    (document-length normalisation)."""

    def __init__(self, texts, k1=1.2, b=0.75):
        self.k1, self.b = k1, b
        self.docs = [Counter(tokens(t)) for t in texts]
        self.lengths = [sum(d.values()) for d in self.docs]
        self.avg_len = sum(self.lengths) / len(self.docs)
        n = len(self.docs)
        df = Counter(term for d in self.docs for term in d)
        # Rare terms weigh more. This IDF form is always positive.
        self.idf = {term: math.log(1 + (n - f + 0.5) / (f + 0.5)) for term, f in df.items()}

    def score(self, query, i):
        doc, length, total = self.docs[i], self.lengths[i], 0.0
        for term in set(tokens(query)):
            tf = doc.get(term, 0)
            if tf:
                norm = tf * (self.k1 + 1) / (tf + self.k1 * (1 - self.b + self.b * length / self.avg_len))
                total += self.idf[term] * norm
        return total


# ---------------------------------------------------------------------------
# Vector search: the Unit 3 model, or toy "concept" embeddings with --offline
# ---------------------------------------------------------------------------
# The toy maps words (English and German) to a few concepts, so "stuck" and "blocked" land in
# the same place. Unknown words, such as codes like PRC-112, are ignored, which is roughly how
# embeddings blur rare identifiers. It is a teaching stand-in, not a real model.
CONCEPTS = {
    "block": "stuck blocked block blocks held hold frozen gesperrt sperre",
    "credit": "credit exposure limit rating kreditlimit obligo kreditanalyst",
    "release": "release released releases approve approves approved approval unblock freigabe gibt frei",
    "sales": "customer customers buyer client sales kunde kundenauftrag",
    "order": "order orders auftrag kundenauftrag",
    "invoice": "invoice invoices billed bill billing charged",
    "receipt": "receipt delivered delivery goods pieces quantity fewer",
    "variance": "differs difference variance mismatch more higher tolerance agreed",
    "purchase": "purchase vendor supplier procurement buyer",
    "price": "price pricing prices condition conditions discount",
    "error": "error errors problem missing mandatory mean fix",
    "shortage": "shortage missing short available lacking expedited",
    "parts": "component components parts part material assembly",
    "plan": "planning planner plan mrp production reschedule run",
    "pay": "payment pay paid unpaid due terms",
    "date": "date dates week days baseline when",
    "policy": "policy percent manager analyst approves",
}
CONCEPT_OF = {}
for position, (concept, words) in enumerate(CONCEPTS.items()):
    for word in words.split():
        CONCEPT_OF.setdefault(word, []).append(position)


def toy_embed(texts):
    vectors = np.zeros((len(texts), len(CONCEPTS)), dtype=np.float32)
    for row, text in enumerate(texts):
        for token in re.findall(r"\w+", text.lower()):
            for position in CONCEPT_OF.get(token, []):
                vectors[row, position] += 1.0
    norms = np.linalg.norm(vectors, axis=1, keepdims=True)
    return vectors / np.where(norms == 0, 1, norms)


def get_embedder(offline):
    if offline:
        return "toy concept embeddings (--offline)", toy_embed
    try:
        from sentence_transformers import SentenceTransformer
    except ImportError:
        sys.exit("sentence-transformers is not installed. See Set up for Unit 3, or add --offline.")
    try:
        model = SentenceTransformer(MODEL)
    except Exception as error:  # network blocked, proxy, disk full
        sys.exit(f"Could not load the model ({type(error).__name__}). Check your network, or add --offline.")
    return MODEL, lambda texts: model.encode(texts, normalize_embeddings=True)


# ---------------------------------------------------------------------------
# Metadata filters: access first, then precision
# ---------------------------------------------------------------------------
def build_filter(user, doc_type=None, language=None, as_of=None):
    """Return a function that says whether a note may be searched.

    The company code rule comes from USERS (the server's view of the user), never from the
    question. A note without a company code is excluded: when in doubt, leave it out."""
    if user not in USERS:
        sys.exit(f"Unknown user '{user}'. Choose one of: {', '.join(USERS)}")
    allowed = set(USERS[user])

    def passes(note):
        code = note.get("company_code")
        if code is None:                       # missing label: fail closed
            return False
        if code != "ALL" and code not in allowed:
            return False
        if doc_type and note["doc_type"] != doc_type:
            return False
        if language and note["language"] != language:
            return False
        if as_of and not (note["valid_from"] <= as_of <= note["valid_to"]):
            return False
        return True

    return passes


# ---------------------------------------------------------------------------
# Ranking and fusion
# ---------------------------------------------------------------------------
def ranked(scores):
    """(index, score) pairs, best first, dropping zero keyword scores (no word in common)."""
    return sorted(((i, s) for i, s in scores.items() if s > 0), key=lambda pair: -pair[1])


def rrf(lists, k=60, window=20):
    """Reciprocal rank fusion: each list adds 1 / (k + rank) for every note it ranks."""
    fused = {}
    for result in lists:
        for rank, (i, _) in enumerate(result[:window], start=1):
            fused[i] = fused.get(i, 0.0) + 1.0 / (k + rank)
    return sorted(fused.items(), key=lambda pair: -pair[1])


def weighted(keyword, vector, alpha=0.5, window=20):
    """Min-max scale each list's scores to 0..1, then alpha * vector + (1 - alpha) * keyword."""
    def scaled(result):
        result = result[:window]
        if not result:
            return {}
        high, low = result[0][1], result[-1][1]
        if high == low:  # all scores tie: 1.0 if they are real matches, 0.0 if all are zero
            return {i: 1.0 if high > 0 else 0.0 for i, _ in result}
        return {i: (s - low) / (high - low) for i, s in result}

    kw, vec = scaled(keyword), scaled(vector)
    fused = {i: alpha * vec.get(i, 0.0) + (1 - alpha) * kw.get(i, 0.0) for i in set(kw) | set(vec)}
    return sorted(fused.items(), key=lambda pair: -pair[1])


class Searcher:
    def __init__(self, offline):
        self.model_name, self.embed = get_embedder(offline)
        texts = [note["text"] for note in NOTES]
        self.bm25 = BM25(texts)
        self.vectors = np.asarray(self.embed(texts), dtype=np.float32)

    def run(self, question, passes, k=60, alpha=0.5):
        rows = [i for i, note in enumerate(NOTES) if passes(note)]
        query_vector = np.asarray(self.embed([question]), dtype=np.float32)[0]
        keyword = ranked({i: self.bm25.score(question, i) for i in rows})
        vector = sorted(((i, float(self.vectors[i] @ query_vector)) for i in rows), key=lambda p: -p[1])
        return {
            "bm25": keyword,
            "vector": vector,
            "rrf": rrf([keyword, vector], k=k),
            "weighted": weighted(keyword, vector, alpha=alpha),
        }, len(rows)


def cmd_search(args):
    searcher = Searcher(args.offline)
    as_of = args.as_of or None
    passes = build_filter(args.user, args.doc_type, args.lang, as_of)
    results, searched = searcher.run(args.question, passes, k=args.k, alpha=args.alpha)
    print(f"Embeddings: {searcher.model_name}")
    print(f"User {args.user} may see company codes {', '.join(USERS[args.user])} (plus ALL)")
    print(f"Searched {searched} of {len(NOTES)} notes after filters\n")
    modes = ["bm25", "vector", "rrf", "weighted"] if args.mode == "all" else [args.mode]
    for mode in modes:
        print(f"[{mode}]")
        if not results[mode]:
            print("  no matching notes (a valid empty result: nothing passed the filters or matched)")
        for rank, (i, score) in enumerate(results[mode][:args.top], start=1):
            note = NOTES[i]
            print(f"  {rank}. {note['id']} [{note.get('company_code', '?')}, {note['doc_type']}, "
                  f"{note['language']}] {score:.3f}  {note['text'][:70]}...")
        print()


def cmd_eval(args):
    searcher = Searcher(args.offline)
    modes = ["bm25", "vector", "rrf", "weighted"]
    hits = {m: 0 for m in modes}
    reciprocal = {m: 0.0 for m in modes}
    print(f"Embeddings: {searcher.model_name}\n")
    print(f"{'question':<52}" + "".join(f"{m:>9}" for m in modes))
    for question, user, expected in EVAL_SET:
        results, _ = searcher.run(question, build_filter(user), k=args.k, alpha=args.alpha)
        row = f"{question[:50]:<52}"
        for mode in modes:
            ids = [NOTES[i]["id"] for i, _ in results[mode]]
            rank = ids.index(expected) + 1 if expected in ids else None
            hits[mode] += 1 if rank and rank <= 3 else 0
            reciprocal[mode] += 1 / rank if rank else 0.0
            row += f"{(str(rank) if rank else '-'):>9}"
        print(row)
    n = len(EVAL_SET)
    print("\nRank of the expected note per method ('-' = not found)")
    print(f"{'hit@3':<52}" + "".join(f"{hits[m] / n:>9.2f}" for m in modes))
    print(f"{'MRR':<52}" + "".join(f"{reciprocal[m] / n:>9.2f}" for m in modes))
    print(f"\nrrf k={args.k}, weighted alpha={args.alpha} (1.0 = vector only, 0.0 = keyword only)")


def valid_date(text):
    try:
        return date.fromisoformat(text).isoformat()
    except ValueError:
        raise argparse.ArgumentTypeError("use the form YYYY-MM-DD, for example 2026-10-01")


def main():
    parser = argparse.ArgumentParser(description="Hybrid search with metadata filters.")
    sub = parser.add_subparsers(dest="command", required=True)
    search = sub.add_parser("search", help="search the notes")
    search.add_argument("question")
    search.add_argument("--user", default="ben", help="anna, ben or chen (default ben)")
    search.add_argument("--mode", default="rrf", choices=["bm25", "vector", "rrf", "weighted", "all"])
    search.add_argument("--doc-type", choices=["policy", "case_note", "help"])
    search.add_argument("--lang", choices=["en", "de"])
    search.add_argument("--as-of", type=valid_date, help="only notes valid on this date, YYYY-MM-DD")
    search.add_argument("--top", type=int, default=3)
    evaluate = sub.add_parser("eval", help="score every method on the test questions")
    for p in (search, evaluate):
        p.add_argument("--offline", action="store_true", help="toy concept embeddings; no model needed")
        p.add_argument("--k", type=int, default=60, help="RRF constant (default 60)")
        p.add_argument("--alpha", type=float, default=0.5, help="weight of vector scores (default 0.5)")
    args = parser.parse_args()
    if not 0.0 <= args.alpha <= 1.0:
        sys.exit("--alpha must be between 0.0 and 1.0")
    {"search": cmd_search, "eval": cmd_eval}[args.command](args)


if __name__ == "__main__":
    main()

Step 3: See each ranker's blind spot

  1. Search for a material number and show every method side by side:

    python unit07/hybrid_search.py search "RM-4711" --mode all --offline --top 2
  2. You should see:

    Embeddings: toy concept embeddings (--offline)
    User ben may see company codes 1000, 2000 (plus ALL)
    Searched 14 of 16 notes after filters
    
    [bm25]
      1. N10 [1000, case_note, en] 2.444  Planning run showed a shortage of component RM-4711 for production in ...
    
    [vector]
      1. N01 [1000, policy, en] 0.000  Credit policy v1: a sales order blocked by the credit check may be rel...
      2. N02 [1000, policy, en] 0.000  Credit policy v2: a sales order blocked by the credit check may be rel...
    
    [rrf]
      1. N10 [1000, case_note, en] 0.031  Planning run showed a shortage of component RM-4711 for production in ...
      2. N01 [1000, policy, en] 0.016  Credit policy v1: a sales order blocked by the credit check may be rel...
    
    [weighted]
      1. N10 [1000, case_note, en] 0.500  Planning run showed a shortage of component RM-4711 for production in ...
      2. N01 [1000, policy, en] 0.000  Credit policy v1: a sales order blocked by the credit check may be rel...

    BM25 found the one note with that code and nothing else. Vector search found nothing meaningful, but it still returned notes, all with similarity 0.000. Vector search always returns something. Both fusion methods put the right note first.

  3. Now ask in plain words, with no word from the right note:

    python unit07/hybrid_search.py search "lacking components to build product" --mode all --offline --top 2

    You should see (header lines and the [weighted] block left out here):

    [bm25]
      1. N11 [ALL, help, en] 2.319  MRP exceptions: the planning run flags missing parts, late receipts an...
    
    [vector]
      1. N12 [2000, case_note, en] 0.686  Missing parts for assembly line 2: component RM-5120 not delivered. Pl...
      2. N11 [ALL, help, en] 0.442  MRP exceptions: the planning run flags missing parts, late receipts an...
    
    [rrf]
      1. N11 [ALL, help, en] 0.033  MRP exceptions: the planning run flags missing parts, late receipts an...
      2. N12 [2000, case_note, en] 0.016  Missing parts for assembly line 2: component RM-5120 not delivered. Pl...

    The case note N12 is the best match, and only vector search ranks it first. BM25 matched only the word "build" in the help note N11. RRF puts N11 first because both lists contain it. This is RRF's known trade-off: agreement beats a single strong hit. Keep that in mind for Step 5.

Step 4: Add filters, for precision and for access

  1. Ask as Anna, who may see only company code 1000:

    python unit07/hybrid_search.py search "credit limit release" --user anna --offline
    User anna may see company codes 1000 (plus ALL)
    Searched 11 of 16 notes after filters
    
    [rrf]
      1. N01 [1000, policy, en] 0.033  Credit policy v1: a sales order blocked by the credit check may be rel...
      2. N02 [1000, policy, en] 0.032  Credit policy v2: a sales order blocked by the credit check may be rel...
      3. N03 [1000, case_note, en] 0.031  Customer order stuck for three days. Cause: credit exposure above the ...

    The access filter already removed five notes: those of company codes 2000 and 3000, and N16, which has no company code at all. But the top result is policy v1, which expired on 30 June 2026.

  2. Add precision filters: only notes valid on 1 October 2026, in English:

    python unit07/hybrid_search.py search "credit limit release" --user anna --as-of 2026-10-01 --lang en --offline
    Searched 9 of 16 notes after filters
    
    [rrf]
      1. N02 [1000, policy, en] 0.033  Credit policy v2: a sales order blocked by the credit check may be rel...
      2. N03 [1000, case_note, en] 0.032  Customer order stuck for three days. Cause: credit exposure above the ...
      3. N05 [ALL, help, en] 0.016  Error PRC-112 in pricing means a mandatory price condition is missing....

    The current policy is now first. No ranking trick could have done that reliably: v1 and v2 use almost the same words.

  3. Ask the same question as Chen, who may see all three company codes:

    python unit07/hybrid_search.py search "customer order blocked credit limit" --user chen --offline --top 4
    User chen may see company codes 1000, 2000, 3000 (plus ALL)
    Searched 15 of 16 notes after filters
    
    [rrf]
      1. N13 [3000, case_note, en] 0.033  Subsidiary customer order blocked: credit limit exceeded after a merge...
      2. N04 [2000, case_note, en] 0.032  Customer order held because the credit limit was exceeded. The analyst...
      3. N01 [1000, policy, en] 0.031  Credit policy v1: a sales order blocked by the credit check may be rel...
      4. N02 [1000, policy, en] 0.031  Credit policy v2: a sales order blocked by the credit check may be rel...

    N13 from company code 3000 is the best match, and Chen may see it. Anna never would. Note that Chen still searched only 15 notes: N16 has no company code, so it is excluded for everyone.

  4. Try a filter combination that matches nothing:

    python unit07/hybrid_search.py search "Kreditlimit" --user anna --doc-type policy --lang de --offline
    Searched 0 of 16 notes after filters
    
    [rrf]
      no matching notes (a valid empty result: nothing passed the filters or matched)

    An empty result is valid. In an assistant, it should lead to "I found no policy for this", not to a guess.

Step 5: Measure every method

  1. Run the evaluation:

    python unit07/hybrid_search.py eval --offline
  2. You should see:

    Embeddings: toy concept embeddings (--offline)
    
    question                                                 bm25   vector      rrf weighted
    PRC-112                                                     1        5        1        1
    4500017311                                                  1        7        1        1
    RM-4711 status                                              1       10        1        1
    RM-5120                                                     1       12        1        1
    client purchase frozen because they had not paid            -        2        4        3
    supplier sent fewer pieces than we were billed for          1        1        1        1
    vendor charged more than agreed                             1        1        1        1
    lacking components to build product                         -        1        2        2
    who may release a credit hold above 3 percent               1        1        1        1
    what does error PRC-112 mean                                1        1        1        1
    
    Rank of the expected note per method ('-' = not found)
    hit@3                                                    0.80     0.60     0.90     1.00
    MRR                                                      0.80     0.60     0.88     0.88
    
    rrf k=60, weighted alpha=0.5 (1.0 = vector only, 0.0 = keyword only)
  3. Read it:

    • The first four questions are codes. BM25 ranks them first; vector search ranks them 5th to 12th.
    • Two paraphrase questions share no useful word with the right note. BM25 doesn't find them at all (-); vector search does.
    • Both fusion methods beat either ranker alone on hit@3. Neither is perfect.
  4. Change the fusion settings and compare:

    python unit07/hybrid_search.py eval --offline --alpha 0.8
    python unit07/hybrid_search.py eval --offline --alpha 1.0
    python unit07/hybrid_search.py eval --offline --k 2

    With --alpha 1.0, weighted fusion is vector search alone, and its MRR drops to 0.60. With --alpha 0.8, MRR rose to 0.93 in our run. --k 2 made no difference on this small set. Ten questions are too few to choose settings for production; they are enough to see the trade-offs. Unit 8 covers building a proper evaluation set.

Step 6: Save your work

  1. Check what Git will save:

    git status

    You should see unit07/hybrid_search.py. The script writes no other files.

  2. Save it:

    git add unit07/hybrid_search.py
    git commit -m "Add hybrid search lab with metadata filters"

What the code does

Part What it does
NOTES 16 made-up notes with metadata. ALL marks shared content; N16 deliberately lacks a company code
USERS Which company codes each user may see. Stands in for a lookup of the user's SAP roles on the server
EVAL_SET 10 test questions, each with a user and the note that should rank first
tokens Lowercases text, keeps codes like prc-112 as one token, drops common words
BM25 Computes IDF per word once, then scores a note with k1 = 1.2 and b = 0.75
toy_embed, CONCEPTS The --offline stand-in: maps words to 17 concepts; unknown words and codes are ignored
get_embedder Loads all-MiniLM-L6-v2 from Unit 3, or the toy with --offline
build_filter Builds the filter: access by company code (missing label fails closed), then document type, language and validity date
ranked Sorts keyword scores and drops notes with score 0 (no shared word)
rrf Adds 1 / (k + rank) from each list for its top 20
weighted Rescales each list to 0..1 (ties become 1.0 for real matches, 0.0 for no signal), then alpha × vector + (1 − alpha) × keyword
Searcher.run Filters first, then ranks only the allowed notes with both methods and fuses
cmd_search, cmd_eval Print results, or hit@3 and MRR per method

If something goes wrong

What you see What it means What to do
python is not recognized / command not found Python isn't on your PATH, or the terminal was open before you installed it Close and reopen the terminal; see Set up your computer for this course
numpy is not installed The library is missing in the Python you're using Check for (.venv) in the prompt, then pip install -r requirements.txt
sentence-transformers is not installed Unit 3's library is missing from this .venv Follow Set up for Unit 3, or add --offline
Could not load the model (...) The model download was blocked, often by a company proxy or firewall Try another network, or ask IT to allow Hugging Face downloads. Meanwhile use --offline
pip install fails with a proxy or SSL error The company network blocks the package index Ask IT to allow pypi.org, or try on another network
Unknown user 'zed' --user must be one of the names in USERS Use anna, ben or chen, or add a user to USERS
argument --as-of: use the form YYYY-MM-DD The date is not a real date in that format For example --as-of 2026-10-01
error: argument command: invalid choice The subcommand is misspelled Use search or eval
no matching notes Nothing passed the filters, or nothing matched A valid result. Remove one filter at a time to see which one empties it
Your rankings differ from the samples You are using the real model, not --offline Expected. Compare the patterns, not the numbers

The SAP way

As of October 2026, here is what SAP documents for the two halves of this topic.

SAP AI Core's release notes of 16 February 2026 state that vector search supports advanced filtering, including metadata filtering, and that you can manage metadata for documents, collections and chunks created with the Vector API. The same release added post-processing in the Retrieval API to merge and rank results from several data repositories.

The SAP Cloud SDK for AI (Java) shows a grounding request that filters on document metadata. This is a sketch: it needs an SAP AI Core instance with the generative AI hub and a grounding collection, set up in Unit 5 and SAP generative AI hub and orchestration.

// SKETCH: needs SAP AI Core with orchestration and a vector grounding collection.
// Class and method names from the SAP Cloud SDK for AI (Java) documentation.
var documentMetadata =
    SearchDocumentKeyValueListPair.create()
        .key("company_code")
        .value("1000")
        .addSelectModeItem(SearchSelectOptionEnum.IGNORE_IF_KEY_ABSENT);

var filter = DocumentGroundingFilter.create()
    .id("")
    .dataRepositoryType(DataRepositoryType.VECTOR)
    .addDocumentMetadataItem(documentMetadata);

Things to plan for:

  • Attach metadata at upload. Company code, document type, language and validity dates go on the document (or chunk) when you add it to a collection. The chunks.jsonl from Chunking and document preparation already carries language and valid_from; add company code and the other access labels the same way.
  • Keyword ranking. This course found no SAP statement, as of October 2026, that the grounding service offers BM25 or hybrid ranking. If your questions are full of codes, test them early, and consider a keyword half outside the service fused in your application.
  • Validity dates. Check which comparisons your release supports. If only equality filters are available, a common workaround is to remove superseded documents from the collection at the moment they expire.

SAP HANA Cloud: filters in SQL, fusion in your code

In SAP HANA Cloud the vector sits in an ordinary table, so filters are ordinary SQL. SAP's learning material shows COSINE_SIMILARITY both as a ranked column and inside a WHERE condition. Adding business filters is the same pattern:

-- SKETCH: example table HELP_NOTES with a REAL_VECTOR column EMBEDDING and label columns.
-- The question vector is passed in as a parameter.
SELECT TOP 20 "ID", "TEXT",
       COSINE_SIMILARITY("EMBEDDING", TO_REAL_VECTOR(?)) AS "SCORE"
  FROM HELP_NOTES
 WHERE "COMPANY_CODE" IN ('1000', 'ALL')
   AND "LANGUAGE" = 'en'
   AND CURRENT_DATE BETWEEN "VALID_FROM" AND "VALID_TO"
 ORDER BY "SCORE" DESC;

SAP's langchain-hana package wraps the same idea in Python. Its HanaDB vector store accepts the filter operators shown earlier, and its specific_metadata_columns option stores chosen metadata keys in their own columns, which its documentation says makes filtering on those keys faster than reading them from the JSON metadata column. Put access labels such as company code in such columns. The package also offers a $contains keyword filter; that requires a word, it doesn't rank by it.

For the keyword half, SAP HANA Cloud has its own text search features. This course has not yet verified, from SAP documentation, how they score results, so the sketch below does the fusion in Python, exactly as the lab does:

# SKETCH: needs an SAP HANA Cloud connection (Set up for Unit 7) and a keyword source.
# keyword_ids and vector_ids are lists of document IDs, best first, from two filtered queries
# that used the SAME company-code and validity conditions.
def rrf_ids(lists, k=60, window=20):
    fused = {}
    for ids in lists:
        for rank, doc_id in enumerate(ids[:window], start=1):
            fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + rank)
    return sorted(fused, key=fused.get, reverse=True)

top_ids = rrf_ids([keyword_ids, vector_ids])[:5]

The SAP HANA Cloud vector engine topic later in Unit 7 goes deeper into the database side.

Licensing notes

  • The lab uses only open-source parts (Python, numpy, optionally sentence-transformers).
  • SAP AI Core grounding usage is billed against your SAP BTP account; check which metrics your contract counts.
  • In SAP HANA Cloud, extra label columns and indexes add memory; include them in sizing. The trial is free but stops nightly.

Build vs. SAP

Need Your own code (this lab) Search engine with hybrid built in SAP HANA Cloud SAP AI Core grounding
Keyword ranking BM25 in a few lines; fine for small sets Mature analyzers and BM25 Own text search; verify scoring for your release Not documented, as of October 2026
Vector ranking numpy or Chroma Built in Vector engine with COSINE_SIMILARITY Built in
Fusion RRF or weighted, your choice Built in; check defaults (k, alpha) In your application code Merging across repositories via Retrieval API post-processing
Metadata filters Any rule you write Built in SQL WHERE on business columns Document and chunk metadata filters since February 2026
Access filters from SAP roles You build the lookup You build the lookup You build the lookup; data often already in SAP You build the lookup
Operations Yours Another system to run Your HANA Cloud team SAP

A practical path: prototype with this lab's approach, write an evaluation set with codes and paraphrases, then pick the store. If many questions contain codes and your chosen SAP service doesn't rank by keywords, fuse in your application. Whatever you pick, the access-filter lookup from SAP roles is yours to build.

Production concerns

  • Authorizations. Derive the access filter on the server from the user's identity and SAP roles. Never take filter values from the browser or the model. Apply the same filter to every ranker.
  • Fail closed. A chunk without an access label is excluded. Report missing labels from ingestion as data errors.
  • Label quality. Labels come from the source document at ingestion. Test that a document's company code reaches all its chunks, and that changes in the source reach the index.
  • Superseded content. Store validity dates, filter on them, and delete expired documents where the store can't filter by date.
  • Evaluation. Keep a test set with three kinds of questions: codes and numbers, paraphrases, and mixed. Track hit@k and MRR per method. Re-run after every change to tokenization, model, fusion settings or data.
  • Thresholds. Vector search returns something for any question. Use the keyword half's "no match" signal or a similarity floor, measured on your data, before telling the model "here is the evidence".
  • Latency. Hybrid runs two searches; run them in parallel. Measure p95 latency with filters on, because filters can speed up or slow down a search depending on the engine.
  • Logging and audit. Log user, filter applied, and returned chunk IDs. OWASP recommends detailed retrieval logs to spot suspicious patterns.
  • Clean core. Keep the index side by side on SAP BTP or in SAP HANA Cloud. Read SAP data and labels through released APIs; don't write search data into S/4HANA tables.

Pitfalls

  • Filtering only one ranker. Fusion pulls the forbidden document back from the unfiltered list.
  • Filtering after fusion. You may end up with too few results, and the ranker still read restricted text.
  • Relaxing filters on empty results. Relaxing language is a choice; relaxing company code is a leak.
  • Trusting alpha or k defaults. Products differ: Elasticsearch defaults RRF to k = 60, Qdrant to k = 2; Weaviate's alpha defaults to 0.75.
  • Min-max scaling without edge cases. One result, or all-equal scores, divides by zero or turns noise into 1.0. The lab hit this bug while being written.
  • Splitting codes in tokenization. RM-4711 split into rm and 4711 matches the wrong things.
  • Choosing settings on ten questions. Small sets show trade-offs, not optimal values.

Exercise

Extend the evaluation with your own questions and save a report. The report, unit07/hybrid_eval_report.txt, becomes part of your evaluation set in Unit 8.

  1. Open unit07/hybrid_search.py and find EVAL_SET.

  2. Add four questions of your own at the end of the list, each as ("question", "user", "note id"):

    • one with an exact code or number from a note,
    • one paraphrase that shares no important word with its note,
    • one that only a user with company code 2000 or 3000 should get right (use ben or chen),
    • one about the credit policy that only the current policy (N02) should answer.
  3. Save, then run the evaluation three times into one report.

    Windows (PowerShell):

    python unit07/hybrid_search.py eval --offline | Out-File -Encoding utf8 unit07/hybrid_eval_report.txt
    python unit07/hybrid_search.py eval --offline --alpha 0.8 | Out-File -Append -Encoding utf8 unit07/hybrid_eval_report.txt
    python unit07/hybrid_search.py eval --offline --k 2 | Out-File -Append -Encoding utf8 unit07/hybrid_eval_report.txt

    macOS/Linux:

    python unit07/hybrid_search.py eval --offline > unit07/hybrid_eval_report.txt
    python unit07/hybrid_search.py eval --offline --alpha 0.8 >> unit07/hybrid_eval_report.txt
    python unit07/hybrid_search.py eval --offline --k 2 >> unit07/hybrid_eval_report.txt

    If you have the Unit 3 model, run the same three commands without --offline and append them too.

  4. eval applies only the access filter, not the date filter. Check whether any method ranks the expired policy N01 above N02 for your credit policy question, and note which.

  5. Add three lines by hand at the end of the report: which method you would choose; which question each ranker failed; and one sentence on what filter the credit policy question needs.

  6. Commit:

    git add unit07/hybrid_search.py unit07/hybrid_eval_report.txt
    git commit -m "Add hybrid search evaluation report"

Done when unit07/hybrid_eval_report.txt holds three eval tables with 14 questions each and your three hand-written lines, and EVAL_SET in your script contains your four new questions.

Check yourself

Pick one answer for each question. The explanation appears after you choose.
  1. 1Why does the lab apply the filters before either ranker scores a note?

    Answer: B. Filtering before ranking keeps restricted notes out of both rankers, so fusion can't pull them back. Filtering after ranking can also return fewer results than asked.
  2. 2In the lab, the query "RM-4711" gets vector scores of 0.000 for every note. What does that show?

    Answer: A. With --offline, the toy ignores unknown words, so the code gives an empty meaning and all similarities tie at zero. Real vector search also always returns something, which is why you need a no-match signal or a measured threshold.
  3. 3With RRF and k = 60, why can a note ranked first by only one list lose to a note ranked first and second by both lists?

    Answer: D. Each list adds 1 / (k + rank); with k = 60, rank 1 and rank 2 earn almost the same. A note in both lists collects two contributions and beats a single first place, as with N11 and N12 in the lab.
  4. 4What problem does the lab's weighted function guard against when all scores in a list are equal?

    Answer: B. Scaling uses (score − low) / (high − low), which breaks when high equals low. The lab gives real matches 1.0 and all-zero lists 0.0, so a meaningless vector list doesn't outvote the keyword hit.
  5. 5You add a hybrid search to an SAP AI Core grounding setup and filter company code with select mode IGNORE_IF_KEY_ABSENT. What would you do before go-live?

    Answer: C. Read literally, that select mode does not exclude documents missing the key, which fails open for an access filter. Make labels complete and test the missing-key case on your release.
  6. 6Why should the access filter values come from a server-side lookup of the user's roles?

    Answer: D. Anything the browser or the model sends can be changed. Deriving company codes from the authenticated identity on the server is what turns a filter into an access control.
  7. 7In SAP HANA Cloud, how does the topic suggest combining keyword and vector results?

    Answer: C. The course has not verified how HANA's text search scores results, so the sketch runs both filtered queries and fuses the ID lists in the application. $contains requires a word but does not rank by it.
  8. 8Your evaluation shows BM25 hit@3 of 0.80, vector 0.60 and RRF 0.90 on ten questions. What is the right conclusion?

    Answer: B. Ten questions show the trade-offs, not reliable optimal values. Each ranker failed different questions, so hybrid is the likely choice, but settings need a larger evaluation set, which Unit 8 builds.

Sources

Sign in to track your progress

We'll email you a one-time sign-in link. No password needed.

or

Tell us a little about you

Optional, every field. It helps us pitch answers to your questions at the right level and decide which topics to write next. It is never shown publicly, and you can change or clear it anytime from the account menu.

SAP areas you work in