Krzysztof Duda

Seeing What Search Can't Show: Visualizing 2000 Bookmarks

I have 2000 bookmarks. I've read maybe 50.

They sit in folders like "Read Later" and "Interesting" — digital hoarding disguised as productivity. The problem isn't saving links. It's finding them again.

Search doesn't help. You can't search for something you've forgotten exists.

The Search Problem

Search assumes you know what you're looking for. Type "machine learning tutorial" and you'll find it — if you remember saving it, if you remember what you called it, if your past self used the same words you're thinking now.

But discovery isn't search. Discovery is finding the article you saved three years ago that's suddenly relevant. It's noticing that half your bookmarks are about the same topic you never consciously tracked. It's seeing gaps in your knowledge you didn't know existed.

A list of 2000 items can't show you this. Neither can folders. You need to see the shape of your collection.

Meaning Has Geography

Modern embedding models turn text into coordinates. Feed an article to OpenAI's embedding API and you get back 1536 numbers — a vector that captures semantic meaning.

Similar content produces similar vectors. An article about React hooks and another about Vue composables end up close together in that 1536-dimensional space. Not because they share keywords, but because they're about the same underlying concept: reactive state in frontend frameworks.

The problem: 1536 dimensions don't fit on a screen. PCA (Principal Component Analysis) reduces those dimensions to three by finding the axes where data varies most. It's a linear projection — some nuance gets lost — but neighbors in high-dimensional space tend to stay neighbors in 3D. Good enough for seeing the shape of things.

Now meaning has geography. Your machine learning articles cluster on one side. Your cooking recipes on another. TypeScript and Rust sit near each other — both typed languages. The visualization shows what a list hides: the structure of your interests.

Four Ways to See Your Data

I built four visualization modes, each revealing different patterns.

Tags mode shows your explicit organization. Bookmarks connect to their tags with lines. You see which tags overlap, which bookmarks got multiple labels, where your tagging system breaks down. It's a mirror of how you think you organized things.

loading...

Embeddings mode shows AI-detected topics. The system groups bookmarks by semantic similarity and colors each cluster. You didn't create these groups — the embedding model found them. Often they reveal categories you never explicitly named but always implicitly saved.

loading...

Positions mode uses pre-computed PCA coordinates. No physics simulation, just the raw embedding positions scaled to 3D space. It loads instantly and shows the "true" distances between your bookmarks. Fast and honest.

loading...

Sphere mode projects everything onto a globe. Dense areas become continents — visible concentrations where your interests cluster. You can see at a glance where you've saved the most. The topology reveals your intellectual geography.

loading...

What Discovery Reveals

Users report three types of insights.

Hidden clusters. "I didn't realize I'd saved so many articles about database optimization." The spatial view makes patterns obvious that were invisible in a list.

Unexpected connections. Articles about woodworking sit near articles about software architecture. Both discuss craftsmanship, iteration, building things right. The embedding model found a connection the user never searched for.

Knowledge gaps. A dense cluster about frontend frameworks with nothing about testing. A region about startups with no marketing content. The empty spaces reveal what you're not learning.

None of this shows up in search results. Search gives you what you asked for. Visualization shows you what you didn't know to ask.

You Can Build This

The stack isn't exotic: PostgreSQL with pgvector for storing and searching embeddings, OpenAI's text-embedding-3-small for generating them, Three.js for 3D rendering, d3-force-3d for physics simulation. PCA runs server-side in Ruby — no ML libraries needed, just linear algebra.

The hard part isn't the technology. It's deciding what to show. Four visualization modes exist because no single view captures everything. Tags show your intent. Embeddings show the content. Positions show true distances. Sphere shows density.

The code isn't magic. It's geometry applied to meaning.

Try It

If you have bookmarks scattered across browsers, Pocket, Raindrop, or plain text files — import them into Indeks and see what your interests actually look like.