Text and AI
Semantic search and the cost of dimension reduction
Build a vector search tool and measure accuracy lost as dimensions shrink.
- Status
- Planned
- Dataset
- 20 Newsgroups or arXiv abstracts
- Concepts
- embeddings, cosine similarity, random projection, FAISS
Problem
What question does this project answer, and who would use the answer?
Data
Source, size, licence, and any cleaning applied.
Method
Steps taken and why each was chosen.
Results
Key numbers and charts, with an honest evaluation.
Limitations
What the result does not show.
Theory note
The core idea behind the method, in plain language.