<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Learning with Yahya]]></title><description><![CDATA[A place where I document what I'm learning, what I'm building, and the lessons I discover along the way — from software development and backend systems to AI an]]></description><link>https://learningwithyahya.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Learning with Yahya</title><link>https://learningwithyahya.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 07 Oct 2026 10:26:51 GMT</lastBuildDate><atom:link href="https://learningwithyahya.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Building Consistent Semantic Search: Updates, Embeddings & Concurrency in BrainZen]]></title><description><![CDATA[When I wrote my previous article, I had reached an important point in BrainZen.
The creation pipeline was working:
Content
   ↓
Semantic text
   ↓
Embedding
   ↓
MongoDB

And I had started designing t]]></description><link>https://learningwithyahya.hashnode.dev/building-consistent-semantic-search-updates-embeddings-concurrency-in-brainzen</link><guid isPermaLink="true">https://learningwithyahya.hashnode.dev/building-consistent-semantic-search-updates-embeddings-concurrency-in-brainzen</guid><category><![CDATA[semantic search]]></category><category><![CDATA[#Embeddings]]></category><category><![CDATA[MongoDB]]></category><category><![CDATA[TypeScript]]></category><category><![CDATA[Backend Development]]></category><dc:creator><![CDATA[Mohammed Yahya Ahmed]]></dc:creator><pubDate>Fri, 25 Sep 2026 19:16:28 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ab190a531115ca0c8cca46b/e5adc9ca-cc8d-47c2-b969-657cb70f6c88.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When I wrote my previous article, I had reached an important point in BrainZen.</p>
<p>The creation pipeline was working:</p>
<pre><code class="language-plaintext">Content
   ↓
Semantic text
   ↓
Embedding
   ↓
MongoDB
</code></pre>
<p>And I had started designing the search system around both semantic and lexical retrieval.</p>
<p>At that point, I thought the next challenge would simply be building the search functionality.</p>
<p>But before search could become reliable, I had to answer a more fundamental question:</p>
<blockquote>
<p><strong>What happens when the content that was embedded changes?</strong></p>
</blockquote>
<p>That question led me into a completely different part of system design.</p>
<p>I had to think about <strong>updates, stale embeddings, partial updates, authorization, versioning, race conditions, and atomic database operations.</strong></p>
<p>The interesting part was that none of these problems existed in the same way during creation.</p>
<hr />
<h1>The Problem With Updating Embedded Content</h1>
<p>Suppose BrainZen stores:</p>
<pre><code class="language-plaintext">Title:
React Performance

Description:
Reducing unnecessary renders

Tags:
react, frontend
</code></pre>
<p>and generates:</p>
<pre><code class="language-plaintext">Embedding:
E1
</code></pre>
<p>The relationship is:</p>
<pre><code class="language-plaintext">Content State 1
      ↕
     E1
</code></pre>
<p>Now imagine the user changes the title:</p>
<pre><code class="language-plaintext">React Performance
        ↓
React Performance Optimization
</code></pre>
<p>The content has changed.</p>
<p>But if the database still contains:</p>
<pre><code class="language-plaintext">Embedding = E1
</code></pre>
<p>then we have:</p>
<pre><code class="language-plaintext">New Content
      ↕
   Old E1
</code></pre>
<p>The database technically contains valid data, but the semantic representation is now stale.</p>
<p>That means a future semantic search could retrieve the document based on an embedding that represents the <strong>previous version of the content</strong>.</p>
<p>So updating content isn't simply:</p>
<pre><code class="language-plaintext">change a field
</code></pre>
<p>It becomes:</p>
<pre><code class="language-plaintext">change content
        +
possibly regenerate embedding
        +
keep both states consistent
</code></pre>
<p>That was the first major realization.</p>
<hr />
<h1>Not Every Update Requires a New Embedding</h1>
<p>The next question was:</p>
<blockquote>
<p>Does every update require generating a new embedding?</p>
</blockquote>
<p>No.</p>
<p>BrainZen has different kinds of fields.</p>
<p>For example:</p>
<pre><code class="language-plaintext">title
description
tags
</code></pre>
<p>contribute to the semantic representation.</p>
<p>But:</p>
<pre><code class="language-plaintext">link
type
</code></pre>
<p>can serve different purposes.</p>
<p>So I separated updates into two categories.</p>
<h3>Semantic changes</h3>
<pre><code class="language-plaintext">title
description
tags
</code></pre>
<p>If these change:</p>
<pre><code class="language-plaintext">new content
    ↓
new semantic representation
    ↓
new embedding
</code></pre>
<h3>Non-semantic changes</h3>
<p>For example:</p>
<pre><code class="language-plaintext">link
</code></pre>
<p>If only the link changes:</p>
<pre><code class="language-plaintext">new content
    ↓
same semantic meaning
    ↓
existing embedding can remain
</code></pre>
<p>This gave me an important rule:</p>
<blockquote>
<p><strong>The need for a new embedding should be determined by changes to the semantic representation, not by the fact that an update happened.</strong></p>
</blockquote>
<hr />
<h1>Partial Updates Make This More Interesting</h1>
<p>BrainZen uses a PATCH-style update.</p>
<p>That means the client doesn't necessarily send the entire document.</p>
<p>For example:</p>
<pre><code class="language-plaintext">{
  "title": "React Performance Optimization",
  "version": 5
}
</code></pre>
<p>This means:</p>
<blockquote>
<p>Change the title, but keep everything else.</p>
</blockquote>
<p>So the backend can't simply look at the request and assume that the request represents the complete new document.</p>
<p>It has to construct the <strong>final state</strong>.</p>
<p>Conceptually:</p>
<pre><code class="language-plaintext">Existing document
        +
PATCH fields
        ↓
Final document state
</code></pre>
<p>For example:</p>
<pre><code class="language-plaintext">Existing:
title       = React Performance
description = Reducing renders
tags        = react, frontend

PATCH:
title       = React Performance Optimization

Final:
title       = React Performance Optimization
description = Reducing renders
tags        = react, frontend
</code></pre>
<p>Only after constructing this final state can the backend determine:</p>
<pre><code class="language-plaintext">Did the semantic representation actually change?
</code></pre>
<p>This was an important distinction for me.</p>
<p>The request isn't necessarily the new document.</p>
<p>It is an instruction for <strong>how to construct the new document</strong>.</p>
<hr />
<h1>Tags Made Updates More Complicated</h1>
<p>Tags already had two representations during creation.</p>
<p>The frontend works with:</p>
<pre><code class="language-plaintext">["react", "frontend"]
</code></pre>
<p>while MongoDB stores tag references as ObjectIds.</p>
<p>That becomes important during updates.</p>
<p>Suppose the database contains:</p>
<pre><code class="language-plaintext">ObjectId("A")
ObjectId("B")
</code></pre>
<p>but the frontend sends:</p>
<pre><code class="language-plaintext">["react", "frontend"]
</code></pre>
<p>The backend has to resolve those names again before comparing or persisting them.</p>
<p>And when determining whether tags changed, order shouldn't matter.</p>
<p>These should represent the same tag set:</p>
<pre><code class="language-plaintext">["react", "frontend"]
</code></pre>
<p>and:</p>
<pre><code class="language-plaintext">["frontend", "react"]
</code></pre>
<p>So I normalize the values before comparing them.</p>
<p>The larger lesson was:</p>
<blockquote>
<p><strong>Before comparing structured data, I need to make sure I'm comparing the same representation.</strong></p>
</blockquote>
<p>Otherwise I can accidentally detect a change that isn't really a change.</p>
<hr />
<h1>The Backend Can't Simply Trust the Frontend Embedding</h1>
<p>Because BrainZen currently generates embeddings in the browser, another question appeared:</p>
<blockquote>
<p>How does the backend know that the embedding actually corresponds to the new content?</p>
</blockquote>
<p>The answer is: <strong>it can't completely prove that with the current architecture.</strong></p>
<p>The backend can verify that:</p>
<pre><code class="language-plaintext">embedding exists
        ↓
embedding has 768 dimensions
</code></pre>
<p>But it isn't independently running the embedding model to verify:</p>
<pre><code class="language-plaintext">title + description + tags
        ↓
this exact embedding
</code></pre>
<p>That would require generating the embedding on the backend as well.</p>
<p>So I separated the responsibilities.</p>
<p>The frontend handles:</p>
<pre><code class="language-plaintext">semantic text
     ↓
embedding generation
</code></pre>
<p>The backend handles:</p>
<pre><code class="language-plaintext">validation
authorization
consistency rules
version protection
persistence
</code></pre>
<p>This made the boundary much clearer.</p>
<hr />
<h1>Then I Ran Into Concurrency</h1>
<p>This was probably the most important part of the update design.</p>
<p>Imagine the user changes a document twice very quickly.</p>
<p>We can have:</p>
<pre><code class="language-plaintext">Update A
old content → generate E2

Update B
old content → generate E3
</code></pre>
<p>Both requests started from:</p>
<pre><code class="language-plaintext">version = 5
</code></pre>
<p>Now imagine:</p>
<pre><code class="language-plaintext">E3 finishes first
    ↓
saved successfully

E2 finishes later
    ↓
saved successfully
</code></pre>
<p>Without some protection, the older operation could overwrite the newer one.</p>
<p>We could end up with something like:</p>
<pre><code class="language-plaintext">Content = state from B
Embedding = state from A
</code></pre>
<p>That violates the relationship we were trying to protect.</p>
<p>This is where I learned about <strong>optimistic concurrency control</strong>.</p>
<hr />
<h1>Versioning as a Concurrency Guard</h1>
<p>Instead of assuming that an update is still working with the latest document, the client sends the version it originally read.</p>
<p>For example:</p>
<pre><code class="language-plaintext">Frontend has:
version = 5
</code></pre>
<p>It sends:</p>
<pre><code class="language-plaintext">{
  "title": "New Title",
  "version": 5
}
</code></pre>
<p>The backend then performs an update only if the database still contains:</p>
<pre><code class="language-plaintext">version = 5
</code></pre>
<p>Conceptually:</p>
<pre><code class="language-plaintext">Update request
      ↓
"Update this document IF version = 5"
      ↓
Database
</code></pre>
<p>If the document is still version 5:</p>
<pre><code class="language-plaintext">5 → 6
</code></pre>
<p>The update succeeds.</p>
<p>But if another request already changed it:</p>
<pre><code class="language-plaintext">Database:
version = 6
</code></pre>
<p>then the condition doesn't match.</p>
<p>The update doesn't happen.</p>
<p>The backend can then return:</p>
<pre><code class="language-plaintext">409 Conflict
</code></pre>
<p>The important thing here is that the application doesn't have to lock the document while someone is editing it.</p>
<p>Instead, it detects that the state became stale when the update is committed.</p>
<p>That's why this is called <strong>optimistic concurrency</strong>.</p>
<hr />
<h1>Why $inc Matters</h1>
<p>MongoDB gives us an atomic increment operation:</p>
<pre><code class="language-plaintext">$inc: {
    version: 1
}
</code></pre>
<p>If the document is currently:</p>
<pre><code class="language-plaintext">version = 5
</code></pre>
<p>MongoDB changes it to:</p>
<pre><code class="language-plaintext">version = 6
</code></pre>
<p>as part of the update.</p>
<p>So the update can conceptually be:</p>
<pre><code class="language-plaintext">Find:
    _id = X
    userId = Y
    version = 5

Then:
    set new fields
    increment version
</code></pre>
<p>This is much safer than treating:</p>
<pre><code class="language-plaintext">read version
calculate version + 1
write version
</code></pre>
<p>as separate application-level operations.</p>
<hr />
<h1>Atomic Update: Another Important Lesson</h1>
<p>I also learned that updating related pieces of state should be treated as one logical operation.</p>
<p>For a content update, I don't want:</p>
<pre><code class="language-plaintext">Update title
     ↓
separate operation
     ↓
Update embedding
     ↓
separate operation
     ↓
Update version
</code></pre>
<p>because a failure between those operations could leave the document in an inconsistent state.</p>
<p>Instead, MongoDB can perform the changes together:</p>
<pre><code class="language-plaintext">$set
  title
  description
  tags
  embedding

$inc
  version
</code></pre>
<p>as one update operation.</p>
<p>The desired invariant becomes:</p>
<pre><code class="language-plaintext">Version N
    ↕
Content state N
    ↕
Embedding representing state N
</code></pre>
<p>That's a much stronger mental model than simply thinking:</p>
<blockquote>
<p>"I have a content document with an embedding field."</p>
</blockquote>
<hr />
<h1>What About Delete?</h1>
<p>Delete turned out to have a similar consistency concern.</p>
<p>The embedding itself doesn't need a separate delete operation because it is stored inside the content document.</p>
<p>So:</p>
<pre><code class="language-plaintext">Delete content
      ↓
embedding disappears with it
</code></pre>
<p>But BrainZen also has related share/link documents.</p>
<p>So deleting content can involve:</p>
<pre><code class="language-plaintext">Content
   ↓
related share
</code></pre>
<p>I therefore treated the deletion of the content and its related share cleanup as one database transaction.</p>
<p>The goal is:</p>
<pre><code class="language-plaintext">Delete content
+
Delete related share
        ↓
both succeed
</code></pre>
<p>or:</p>
<pre><code class="language-plaintext">something fails
        ↓
rollback
</code></pre>
<p>rather than leaving behind a partially cleaned-up state.</p>
<hr />
<h1>The Backend Update Pipeline</h1>
<p>After working through all of this, the update flow became much clearer to me.</p>
<pre><code class="language-plaintext">PATCH /content/:id
        ↓
Authentication
        ↓
Validate request
        ↓
Validate Content ID
        ↓
Find content belonging to user
        ↓
Build final document state
        ↓
Did semantic fields change?
      /        \
    NO          YES
    ↓            ↓
keep          require
embedding     new embedding
      \        /
       ↓
Check expected version
        ↓
Atomic MongoDB update
        ↓
Increment version
        ↓
Return new version
</code></pre>
<p>And the key invariant is:</p>
<pre><code class="language-plaintext">Content state N
      ↕
Embedding for state N
      ↕
Version N
</code></pre>
<hr />
<h1>The Bigger Lesson</h1>
<p>What started as:</p>
<blockquote>
<p>"I need to add semantic search."</p>
</blockquote>
<p>has turned into a much broader systems problem.</p>
<p>The embedding itself is actually only one component.</p>
<p>The real system looks more like:</p>
<pre><code class="language-plaintext">                    Content
                       ↓
              Semantic Representation
                       ↓
                    Embedding
                       ↓
                    Storage
                       ↓
              ┌────────┴────────┐
              ↓                 ↓
        Semantic Search    Lexical Search
              ↓                 ↓
         Vector ranking       BM25
              └────────┬────────┘
                       ↓
                      RRF
                       ↓
                   Filtering
                       ↓
                    Results
</code></pre>
<p>But now there's another dimension around the entire system:</p>
<pre><code class="language-plaintext">Create
  ↓
Update
  ↓
Version
  ↓
Consistency
  ↓
Delete
</code></pre>
<p>Search isn't reliable just because retrieval works.</p>
<p>The <strong>data being retrieved has to remain correct over time</strong>.</p>
<p>That's probably the biggest thing I learned from this part of BrainZen:</p>
<blockquote>
<p><strong>Building a search system isn't only about finding relevant information. It's also about making sure the information being searched is represented correctly as the system changes.</strong></p>
</blockquote>
<p>And now that BrainZen can reason about content changes, embeddings, and consistency, the next challenge is the part I initially thought would be the main feature:</p>
<p><strong>actually building the retrieval system — lexical search, semantic search, hybrid retrieval, filtering, and ranking.</strong></p>
]]></content:encoded></item><item><title><![CDATA[From Embeddings to Search Architecture: Building BrainZen’s Retrieval Foundation]]></title><description><![CDATA[When I first started adding semantic search to BrainZen, I thought the difficult part would be generating embeddings and storing them in MongoDB.
It turns out that was only the beginning.
Once the emb]]></description><link>https://learningwithyahya.hashnode.dev/from-embeddings-to-search-architecture-building-brainzen-s-retrieval-foundation</link><guid isPermaLink="true">https://learningwithyahya.hashnode.dev/from-embeddings-to-search-architecture-building-brainzen-s-retrieval-foundation</guid><category><![CDATA[AI]]></category><category><![CDATA[semantic search]]></category><category><![CDATA[hybrid search]]></category><category><![CDATA[architecture]]></category><category><![CDATA[Architecture Design]]></category><category><![CDATA[Vector Search]]></category><category><![CDATA[#Embeddings]]></category><category><![CDATA[Information Retrieval ]]></category><category><![CDATA[MongoDB]]></category><category><![CDATA[TypeScript]]></category><category><![CDATA[backend]]></category><category><![CDATA[System Design]]></category><category><![CDATA[Machine Learning]]></category><dc:creator><![CDATA[Mohammed Yahya Ahmed]]></dc:creator><pubDate>Thu, 24 Sep 2026 09:59:20 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ab190a531115ca0c8cca46b/dfdf2e41-4cb8-4d6a-94e5-252f6e831ae2.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When I first started adding semantic search to BrainZen, I thought the difficult part would be generating embeddings and storing them in MongoDB.</p>
<p>It turns out that was only the beginning.</p>
<p>Once the embedding pipeline was working, I had to understand what actually happens around that vector: what information should be embedded, how the backend should validate it, how authentication fits into the flow, how tags are represented, and eventually how semantic search will work alongside traditional keyword search.</p>
<p>This made me realize that <strong>search isn't a single feature. It's an entire system.</strong></p>
<h2>What Actually Gets Embedded?</h2>
<p>A BrainZen content document contains several different kinds of information:</p>
<pre><code class="language-plaintext">Content
├── title
├── description
├── tags
├── type
├── userId
├── link
└── embedding
</code></pre>
<p>But not all of these fields have the same purpose.</p>
<p>For semantic search, I decided that the meaningful representation should come from:</p>
<pre><code class="language-plaintext">title + description + tags
</code></pre>
<p>Fields like <code>userId</code> and <code>type</code> serve different purposes. <code>userId</code> is about authorization and data isolation, while <code>type</code> can be used as a structured filter.</p>
<p>This distinction is important because the embedding should represent <strong>what the content means</strong>, not every piece of metadata attached to it.</p>
<p>So before generating a vector, BrainZen first converts the structured content into semantic text.</p>
<p>For example:</p>
<pre><code class="language-plaintext">React Performance
Reducing unnecessary renders
react frontend
</code></pre>
<p>That text is then passed to the embedding model.</p>
<p>I intentionally separated these responsibilities:</p>
<pre><code class="language-plaintext">buildEmbeddingText()
        ↓
semantic text
        ↓
generateEmbedding()
        ↓
vector
</code></pre>
<p>The first function answers <strong>what represents the content?</strong></p>
<p>The second answers <strong>how do I turn that representation into a vector?</strong></p>
<hr />
<h2>The Create Flow Is More Than <code>Model.create()</code></h2>
<p>Once the frontend has the embedding, it sends the content to the backend.</p>
<p>The complete flow looks like this:</p>
<pre><code class="language-plaintext">Frontend
   ↓
Generate embedding
   ↓
POST /api/v1/content
   ↓
Authentication middleware
   ↓
Zod validation
   ↓
Resolve tags
   ↓
Get authenticated userId
   ↓
Create MongoDB document
   ↓
Return response
</code></pre>
<p>One of the important things I learned here is that <strong>the backend cannot trust frontend data simply because my own frontend generated it</strong>.</p>
<p>The request is still runtime data coming from a client.</p>
<p>So the backend validates it with Zod before continuing.</p>
<p>Conceptually:</p>
<pre><code class="language-plaintext">req.body
   ↓
untrusted data
   ↓
Zod
   ↓
validated + normalized data
   ↓
business logic
   ↓
database
</code></pre>
<p>This also means I should continue using the validated result rather than going back to the original <code>req.body</code>.</p>
<hr />
<h2>Authentication and User Data Are Different Things</h2>
<p>Another distinction that became much clearer while implementing the flow was where <code>userId</code> should come from.</p>
<p>The frontend does <strong>not</strong> send the user's ID.</p>
<p>Instead:</p>
<pre><code class="language-plaintext">JWT cookie
    ↓
auth middleware
    ↓
verify token
    ↓
extract userId
    ↓
req.userId
</code></pre>
<p>The controller then uses that authenticated identity when creating the document.</p>
<p>This gives me an important security boundary:</p>
<blockquote>
<p><strong>The client provides the content. The server determines the identity.</strong></p>
</blockquote>
<p>If the frontend were allowed to send the <code>userId</code>, a malicious client could simply replace it with another user's ID.</p>
<p>So authentication isn't just about deciding whether someone is logged in. It also establishes <strong>who the request belongs to</strong>.</p>
<hr />
<h2>Tags Have Two Representations</h2>
<p>Tags introduced another interesting transformation.</p>
<p>The frontend naturally works with:</p>
<pre><code class="language-plaintext">["react", "frontend", "performance"]
</code></pre>
<p>But MongoDB stores references to tag documents as ObjectIds.</p>
<p>So the backend performs:</p>
<pre><code class="language-plaintext">"react"
   ↓
find/create tag
   ↓
ObjectId
</code></pre>
<p>for each tag.</p>
<p>The frontend therefore doesn't need to understand MongoDB's internal representation.</p>
<p>It sends:</p>
<pre><code class="language-plaintext">string[]
</code></pre>
<p>and the backend converts that into:</p>
<pre><code class="language-plaintext">ObjectId[]
</code></pre>
<p>before persistence.</p>
<p>This separation keeps the client-facing representation simple while allowing the database to maintain relationships between documents.</p>
<hr />
<h2>Embeddings Have a Contract</h2>
<p>Another thing I had to understand is that an embedding isn't simply "an array of numbers."</p>
<p>BrainZen currently uses:</p>
<pre><code class="language-plaintext">all-mpnet-base-v2
        ↓
768-dimensional vector
</code></pre>
<p>That dimension is part of the system's contract.</p>
<p>The document embedding and the future query embedding must be generated using the same embedding model and configuration.</p>
<p>Otherwise, even if both outputs happen to be arrays of 768 numbers, they aren't necessarily comparable in a meaningful way.</p>
<p>So the architecture is really:</p>
<pre><code class="language-plaintext">Same model
   ├── document → embedding
   └── query    → embedding
</code></pre>
<p>Both vectors live in the same embedding space, which is what makes similarity search meaningful.</p>
<hr />
<h1>Why Semantic Search Isn't Enough</h1>
<p>Once I understood the embedding pipeline, the next question was:</p>
<p><strong>Should BrainZen use semantic search for everything?</strong></p>
<p>No.</p>
<p>Semantic search is excellent at understanding meaning, but sometimes the exact terminology matters.</p>
<p>Consider:</p>
<pre><code class="language-plaintext">MongoDB ObjectId
</code></pre>
<p>If I'm searching for that specific technical term, I don't necessarily want the system to return something merely because it discusses the broader concept of database identifiers.</p>
<p>This is where <strong>lexical search</strong> becomes important.</p>
<p>Semantic search asks:</p>
<blockquote>
<p>"What content means something similar to this query?"</p>
</blockquote>
<p>Lexical search asks:</p>
<blockquote>
<p>"What content contains or strongly matches these terms?"</p>
</blockquote>
<p>They're solving different problems.</p>
<hr />
<h1>BrainZen's Search Architecture</h1>
<p>That led me toward a hybrid architecture:</p>
<pre><code class="language-plaintext">                 Search Query
                      │
             ┌────────┴────────┐
             ↓                 ↓
      Semantic Search    Lexical Search
             │                 │
             ↓                 ↓
       Vector Ranking       BM25
             │                 │
             └────────┬────────┘
                      ↓
                     RRF
                      ↓
               Final Ranking
                      ↓
             Structured Filters
                      ↓
                  Top Results
</code></pre>
<p>The semantic side will generate a query embedding and retrieve semantically similar documents.</p>
<p>The lexical side will use traditional text retrieval and BM25-style ranking.</p>
<p>The two systems produce independently ranked results.</p>
<p>Those rankings can then be combined using <strong>Reciprocal Rank Fusion (RRF)</strong>.</p>
<p>One important thing I learned here is that RRF isn't simply adding the semantic score to the lexical score.</p>
<p>The two systems produce different types of scores, so combining their raw numbers directly isn't necessarily meaningful.</p>
<p>Instead, RRF works from the <strong>rank positions</strong> of documents in each result list.</p>
<hr />
<h1>Tags Can Play Three Roles</h1>
<p>Tags became particularly interesting because they can participate in multiple parts of the architecture.</p>
<p>They can contribute to the semantic representation:</p>
<pre><code class="language-plaintext">title + description + tags
        ↓
embedding
</code></pre>
<p>They can participate in lexical search:</p>
<pre><code class="language-plaintext">tags
 ↓
keyword retrieval
</code></pre>
<p>And they can also be used as a structured filter:</p>
<pre><code class="language-plaintext">tag = "react"
</code></pre>
<p>So the same piece of data can serve three different purposes:</p>
<pre><code class="language-plaintext">Tags
├── Semantic meaning
├── Lexical matching
└── Structured filtering
</code></pre>
<p>These aren't interchangeable operations.</p>
<p>They're different ways of using the same underlying information.</p>
<hr />
<h1>The Bigger Lesson</h1>
<p>The biggest thing I learned wasn't how to generate a 768-dimensional vector.</p>
<p>It was understanding that <strong>search is a pipeline of responsibilities</strong>.</p>
<p>Content starts as structured user input:</p>
<pre><code class="language-plaintext">Content
   ↓
Semantic representation
   ↓
Embedding
   ↓
MongoDB
   ↓
Retrieval
   ↓
Ranking
   ↓
Filtering
   ↓
Results
</code></pre>
<p>And different parts of that pipeline solve different problems.</p>
<ul>
<li><p><strong>Embeddings</strong> represent semantic meaning.  </p>
</li>
<li><p><strong>Vector search</strong> retrieves semantically similar candidates.  </p>
</li>
<li><p><strong>Lexical search</strong> handles terminology and keyword-based retrieval.  </p>
</li>
<li><p><strong>BM25</strong> ranks lexical matches.  </p>
</li>
<li><p><strong>RRF</strong> combines independently ranked retrieval systems.  </p>
</li>
<li><p><strong>Structured filters</strong> enforce constraints such as user ownership, type, or tags.  </p>
</li>
<li><p><strong>Top-K</strong> controls how many results are returned.</p>
</li>
</ul>
<p>What initially looked like "adding AI search" is turning into something much more interesting: <strong>designing an actual retrieval system.</strong></p>
<p>And now that the creation pipeline and embedding foundation are in place, the next problem becomes even more important:</p>
<p><strong>What happens when the user edits the content?</strong></p>
<p>Because if the content changes but its embedding doesn't, BrainZen's search system can become inconsistent.</p>
<p>That's where I'm heading next.</p>
]]></content:encoded></item><item><title><![CDATA[Choosing the Embedding Architecture for BrainZen]]></title><description><![CDATA[When I first started adding semantic search to BrainZen, I figured the actual vector search algorithms would be the biggest headache.
I was wrong.
The very first hurdle was a much more foundational qu]]></description><link>https://learningwithyahya.hashnode.dev/choosing-the-embedding-architecture-for-brainzen</link><guid isPermaLink="true">https://learningwithyahya.hashnode.dev/choosing-the-embedding-architecture-for-brainzen</guid><category><![CDATA[semantic search]]></category><category><![CDATA[search]]></category><category><![CDATA[Vector Search]]></category><category><![CDATA[Information Retrieval ]]></category><category><![CDATA[AI]]></category><category><![CDATA[#Embeddings]]></category><category><![CDATA[hybrid search]]></category><category><![CDATA[nlp]]></category><category><![CDATA[MongoDB Atlas]]></category><category><![CDATA[TransformersJS]]></category><category><![CDATA[huggingface]]></category><category><![CDATA[software architecture]]></category><category><![CDATA[TypeScript]]></category><category><![CDATA[React]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[lexical-search]]></category><category><![CDATA[System Design]]></category><dc:creator><![CDATA[Mohammed Yahya Ahmed]]></dc:creator><pubDate>Tue, 22 Sep 2026 18:26:03 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ab190a531115ca0c8cca46b/aa05bf05-ce7e-4bd2-872b-f62267a4dc5b.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When I first started adding semantic search to BrainZen, I figured the actual vector search algorithms would be the biggest headache.</p>
<p>I was wrong.</p>
<p>The very first hurdle was a much more foundational question: <em>Where exactly is this embedding model going to live?</em> That single decision dictates your hosting costs, security boundaries, performance bottlenecks, and the entire shape of your infrastructure.</p>
<h2>The Current State of BrainZen</h2>
<p>At a high level, BrainZen is split into three main pieces. The crucial detail here is that both the documents you save and the queries you run have to pass through the exact same embedding model. If they come from different embedding spaces, comparing them is mathematically meaningless.</p>
<p>Plaintext</p>
<pre><code class="language-plaintext">Frontend (React + TypeScript)
  │
  │ generates embeddings
  ↓
Embedding Model (all-mpnet-base-v2)
  │
  │ 768-dimensional vector
  ↓
Backend (Express + TypeScript)
  │
  ↓
MongoDB Atlas
  ├── Content
  ├── Tags
  └── Embeddings
</code></pre>
<h2>Where Should the Model Run?</h2>
<p>I had a few distinct architectural paths I could take, each with its own baggage.</p>
<p><strong>External Embedding API</strong> The easiest route. Send the text from the backend to an external API, get a vector back, and save it to Mongo. But this introduces a third-party dependency, API quotas, latency, and usage costs. I didn't want BrainZen's public deployment tethered to my personal API keys.</p>
<p><strong>Run it in the Backend</strong> This keeps the data in-house. But now my lightweight Vercel backend has to lug around an ML model and a Python-esque runtime. For my current deployment setup, that's immediate technical debt.</p>
<p><strong>Separate Inference Service</strong> A dedicated microservice just for handling models. It’s the "enterprise" answer, but it means managing another server, another CI/CD pipeline, and paying for dedicated compute. Overkill.</p>
<p><strong>The Browser (The Winner)</strong> I ultimately decided to run the model directly on the client.</p>
<p>By using <strong>Transformers.js</strong> with the <code>Xenova/all-mpnet-base-v2</code> model, I can generate embeddings right inside the user's browser. It completely eliminates external APIs, API keys, and dedicated ML servers.</p>
<p>The obvious tradeoff? The user’s browser has to download the model and spend CPU cycles running the inference. It’s not the "perfect" architecture for every app, but it was exactly the right compromise for BrainZen's current constraints.</p>
<h2>Solving the Model-Loading Problem</h2>
<p>Once you push ML to the browser, you immediately hit a performance wall. You absolutely cannot do this:</p>
<ul>
<li><p>Generate embedding 1 → Load the model</p>
</li>
<li><p>Generate embedding 2 → Load the model again</p>
</li>
</ul>
<p>The model needs to be loaded into memory exactly once and reused. I didn't need a heavy Singleton class to fix this; a simple module-level cache handles it perfectly:</p>
<p>TypeScript</p>
<pre><code class="language-plaintext">import { pipeline } from "@huggingface/transformers";

let extractor: any = null;

async function getExtractor() {
  if (!extractor) {
    extractor = await pipeline(
      "feature-extraction",
      "Xenova/all-mpnet-base-v2"
    );
  }
  return extractor;
}
</code></pre>
<p>The first time <code>getExtractor()</code> is called, it downloads and caches the model pipeline. Every subsequent keystroke or document upload just reuses that existing instance in memory.</p>
<h2>Turning Text into Vectors</h2>
<p>The actual function doing the heavy lifting is surprisingly tiny:</p>
<p>TypeScript</p>
<pre><code class="language-plaintext">export async function generateEmbedding(text: string): Promise&lt;number[]&gt; {
  const model = await getExtractor();

  const output = await model(text, {
    pooling: "mean",
    normalize: true
  });

  return output.tolist()[0];
}
</code></pre>
<p>Two parameters do the heavy lifting here. Setting <code>pooling: "mean"</code> condenses the model's token-level gibberish into a single, usable representation of the text. Setting <code>normalize: true</code> prepares the resulting vector for the cosine-similarity math I use later on.</p>
<p>Finally, <code>output.tolist()[0]</code> converts the tensor into a standard JavaScript array of 768 floating-point numbers. That array gets shipped off to the database.</p>
<h2>Why Stick with MongoDB?</h2>
<p>Since I was already using MongoDB as BrainZen's source of truth, standing up a dedicated vector database (like Pinecone or Milvus) felt entirely unnecessary.</p>
<p>I just appended an <code>embedding</code> array to my existing schema:</p>
<p>Plaintext</p>
<pre><code class="language-plaintext">Content
 ├── title
 ├── description
 ├── tags
 ├── type
 ├── userId
 └── embedding  &lt;-- 768 numbers live here
</code></pre>
<p>MongoDB Atlas Vector Search hooks right into this field, letting me run semantic queries natively alongside my standard document fetches.</p>
<h2>Why Semantic Search Isn't Enough</h2>
<p>Here is the most practical lesson I learned during this build: <strong>Semantic search cannot fully replace lexical (keyword) search.</strong></p>
<p>If I search for <code>MongoDB ObjectId</code>, I am looking for that exact, specific technical syntax. A purely semantic model might try to be "helpful" and return documents about abstract database identifiers because they mean the same thing conceptually. That's a terrible user experience.</p>
<p>Semantic search is for <em>meaning</em>. Lexical search is for <em>exact terms</em>.</p>
<p>BrainZen's eventual pipeline uses <strong>Hybrid Search</strong>, funneling both semantic and lexical results through a Reciprocal Rank Fusion (RRF) algorithm to score the best possible matches before applying standard filters.</p>
<h2>The Real Takeaway</h2>
<p>The fun part of this wasn't actually writing the <code>generateEmbedding()</code> function. It was figuring out where that computation actually belonged.</p>
<p>Choosing a model isn't just about picking the one with the best benchmark scores. You are inherently choosing your deployment strategy, your infrastructure costs, and how the rest of your system has to communicate.</p>
<p>Instead of just <code>npm install</code>-ing a search library and calling it a day, building BrainZen this way forced me to actually understand why the architecture works in the first place.</p>
]]></content:encoded></item><item><title><![CDATA[Building Search for BrainZen: From Embeddings to Hybrid Retrieval]]></title><description><![CDATA[Search initially seems simple: type a query, get matching content from the database.
But once I started trying to make BrainZen—my second-brain app—actually understand what I'm looking for by meaning ]]></description><link>https://learningwithyahya.hashnode.dev/building-search-for-brainzen-from-embeddings-to-hybrid-retrieval</link><guid isPermaLink="true">https://learningwithyahya.hashnode.dev/building-search-for-brainzen-from-embeddings-to-hybrid-retrieval</guid><category><![CDATA[semantic search]]></category><category><![CDATA[search]]></category><category><![CDATA[Vector Search]]></category><category><![CDATA[Information Retrieval ]]></category><category><![CDATA[AI]]></category><category><![CDATA[#Embeddings]]></category><category><![CDATA[hybrid search]]></category><category><![CDATA[nlp]]></category><category><![CDATA[MongoDB Atlas]]></category><dc:creator><![CDATA[Mohammed Yahya Ahmed]]></dc:creator><pubDate>Mon, 21 Sep 2026 21:01:52 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ab190a531115ca0c8cca46b/24432fe0-61e0-4c3f-97c8-a0b44b71bf74.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Search initially seems simple: type a query, get matching content from the database.</p>
<p>But once I started trying to make BrainZen—my second-brain app—actually understand <em>what</em> I'm looking for by meaning while still catching exact keywords, tags, and names, I realized real search is a whole different beast. Over the past few days, I dove deep into how modern search systems work and started architecting the search layer for BrainZen. Here's a breakdown of what I ran into and what I learned.</p>
<h3>1. Semantic Search (Or: Searching by Meaning)</h3>
<p>Traditional keyword search basically asks one rigid question: <em>"Does this document contain the exact words I typed?"</em></p>
<p>Semantic search, on the other hand, asks: <em>"Is this document actually about the same idea as my query?"</em></p>
<p>To pull this off, you use <strong>embeddings</strong>. An embedding model takes a chunk of text and turns it into a high-dimensional vector of floating-point numbers. For example:</p>
<p>Plaintext</p>
<pre><code class="language-plaintext">"I want to learn how databases organize information" 
↓ (Embedding Model) 
[0.12, -0.84, 0.31, ...]
</code></pre>
<p>The individual numbers don't mean much on their own to a human. The magic is that pieces of text with similar meanings get mapped to vectors that sit close to each other in a multi-dimensional vector space. That lets you compare a query and a document based on conceptual closeness rather than spelling.</p>
<p>To do this efficiently at scale, you rely on <strong>approximate nearest-neighbor (ANN)</strong> algorithms like <strong>HNSW (Hierarchical Navigable Small World)</strong>, which lets you search massive vector spaces quickly without having to do an exhaustive mathematical comparison against every single vector in your database.</p>
<h3>2. Why Semantic Search Isn't Enough</h3>
<p>This was probably the biggest reality check for me.</p>
<p>Suppose I search BrainZen for <code>MongoDB ObjectId</code>. A purely semantic model might deduce that my query has something to do with database identifiers. But sometimes, I don't want a vaguely related philosophical note on IDs—I want the exact string <code>ObjectId</code>.</p>
<p>That’s where <strong>lexical search</strong> comes back into play. Lexical search focuses entirely on exact terms, making it indispensable for:</p>
<ul>
<li><p>Technical terms, names, and exact phrases</p>
</li>
<li><p>Product codes, error messages, and code snippets</p>
</li>
<li><p>Specific tags and identifiers</p>
</li>
</ul>
<p>In short: semantic and lexical search solve two totally different problems. You really want both.</p>
<h3>3. How Lexical Search Works Under the Hood</h3>
<p>On the lexical side, everything revolves around the <strong>inverted index</strong>. Instead of mapping documents to their words, an inverted index flips it around to map terms directly to the documents containing them:</p>
<ul>
<li><p><code>MongoDB</code> \(\rightarrow\) Document 1, Document 3</p>
</li>
<li><p><code>React</code> \(\rightarrow\) Document 2</p>
</li>
<li><p><code>ObjectId</code> \(\rightarrow\) Document 3</p>
</li>
</ul>
<p>Building a proper lexical pipeline means dealing with text processing steps like <strong>tokenization, normalization, stop-word removal, stemming, and lemmatization</strong>, all of which feed into scoring algorithms like <strong>BM25</strong>. BM25 is the industry workhorse for lexical ranking because it doesn't just check if a word exists—it accounts for term frequency, document length, and how rare or common that term is across your entire collection.</p>
<h3>4. Putting Them Together: Hybrid Search &amp; RRF</h3>
<p>So, I don't have to choose. I can run both lexical and semantic retrieval in parallel, but that creates a new issue: <strong>score mismatch</strong>.</p>
<p>A raw vector distance score and a BM25 lexical score speak totally different languages; you can't just add them together mathematically. To solve this, I looked into <strong>Reciprocal Rank Fusion (RRF)</strong>.</p>
<p>Instead of comparing raw scores, RRF looks at the <em>rank position</em> of a document in each result list. If a document pops up near the top of both the lexical search and the semantic search, RRF boosts its final standing. It’s a clean way to combine evidence from multiple systems without overcomplicating the math.</p>
<p>Plaintext</p>
<pre><code class="language-plaintext">               User Query
                   │
         ┌─────────┴─────────┐
         ↓                   ↓
   Lexical Search      Semantic Search
         │                   │
         └─────────┬─────────┘
                   ↓
                  RRF
                   ↓
             Top-K Results
</code></pre>
<h3>5. The Catch: Embedding Lifecycle Management</h3>
<p>Once I mapped out the pipeline, I realized generating embeddings isn't a "set it and forget it" task.</p>
<p>Imagine a note starts like this:</p>
<blockquote>
<p><em>Title:</em> Learning MongoDB</p>
<p><em>Description:</em> Understanding indexes and query optimization</p>
</blockquote>
<p>The system generates an embedding. But what happens tomorrow when I edit the note to add aggregation pipelines? That original vector is suddenly stale.</p>
<p>Building a real app means handling the whole embedding lifecycle: tracking content updates, detecting stale vectors, generating new ones asynchronously, handling API failures, and cleaning up vector storage when a note gets deleted. Keeping your raw content and your vector indices in sync turns out to be a fun engineering challenge on its own.</p>
<h3>Wrapping Up</h3>
<p>Working on BrainZen taught me that search isn't a single algorithm—it's an entire pipeline.</p>
<p>From text preprocessing and inverted indexes to vector spaces, RRF fusion, and metadata filtering, every piece handles a specific failure mode of the others. Lexical finds the exact matches, semantic finds the meaning, RRF merges them, and filters make sure you only see what you're actually allowed to look at.</p>
<p>I'm still actively building and refining this in BrainZen. If there's one thing I've learned, it's that building things from scratch forces you to understand <em>why</em> tools like Elasticsearch, Pinecone, or pgvector do what they do under the hood.</p>
<p>That's what <em>Learning with Yahya</em> is all about: build something, break it, figure out why it broke, and write down what happened.</p>
]]></content:encoded></item></channel></rss>