WEEK3

Embeddings & Vector Search

Semantic search over documents by meaning.

Course Details

Students learn how modern AI search systems understand information using embeddings instead of keywords. They explore document chunking strategies, vector databases, semantic retrieval, and build complete ingestion pipelines that transform raw documents into searchable knowledge. By the end of the week, students can create AI applications that retrieve information based on meaning rather than exact text matches.

Topics Covered

  1. Embeddings & cosine similarity

  2. Chunking strategies & their tradeoffs

  3. Vector databases & indexing

  4. Semantic retrieval by meaning

  5. Building an ingestion pipeline

Topics Covered

OpenAI Embeddings

Pinecone / Qdrant

pgvector

Python

Project 3

KnowledgeVault · PaperFinder

Orange discs and stones balanced with clear glass
Orange discs and stones balanced with clear glass

Project Details

Project 1: Knowledge Vault

Project Description

Build a private AI knowledge base that ingests documents, splits them into meaningful chunks, generates embeddings, and stores them in a vector database. Students implement semantic search to retrieve the most relevant information based on meaning instead of keyword matching.

Project Description

Build an AI-powered research paper search engine that indexes academic papers using embeddings and retrieves the most relevant papers through semantic similarity. Students create a scalable ingestion and indexing pipeline that enables intelligent search across research documents.



Project 1: Knowledge Vault

Project Description

Build a private AI knowledge base that ingests documents, splits them into meaningful chunks, generates embeddings, and stores them in a vector database. Students implement semantic search to retrieve the most relevant information based on meaning instead of keyword matching.

Project Description

Build an AI-powered research paper search engine that indexes academic papers using embeddings and retrieves the most relevant papers through semantic similarity. Students create a scalable ingestion and indexing pipeline that enables intelligent search across research documents.



Project Results

Project Result 1

Developed a document intelligence application that converts unstructured documents into a searchable knowledge base using embeddings, chunking, vector indexing, and semantic retrieval. Implemented an end-to-end ingestion pipeline for efficient AI-powered document search.

Project Result 2

Created a semantic research paper search application using embeddings, vector databases, and cosine similarity. Built a complete ingestion, indexing, and retrieval pipeline that finds relevant papers based on meaning rather than keyword matching.

a planet in space
a planet in space
a group of bottles and glasses
a group of bottles and glasses