This episode discusses the "rise and fall" of the vector database category, arguing that while vector database companies may adapt, the category itself is dying as vector search capabilities are integrated into existing database and search engine technologies. Guest Jokristian Bergam explains that 'search' is a more natural abstraction for Retrieval Augmented Generation (RAG) than 'vector database,' and provides insights into RAG best practices, the long context window debate, and the future of embedding models and knowledge graphs.
Okay. Hi. So this is another Lightning Pod with Jokristian Bergam.
Is that did I get it right? You're over Norway? I'm over in Norway.
Trondheim, Norway, in the center of Norway. Yes. What should people know about Trondheim?
It's a small city. It's easy to get around. There's a great technical university here.
The climate sucks a little bit, but it's easy to get things done in the winter. So yeah, even in Florida. I've never been over.
I think, which is over near you guys. But, yeah, what what we're to talk about just generally your hot takes on Rags, search, vector databases, all that stuff. I think you you've taken to publishing a lot more recently on on X, and that's gone really well.
So I'll just kinda go into that main thing that is that everybody knows you for, which is the, your piece on the vector databases, the rise and fall vector databases. So maybe you could give us the background of, like, why you felt compelled to write this?
Yeah. Of all, I think I have to go a little bit back. Right?
So I have a long background in search and working on infrastructure for search. Like I've been in search, working on search systems for twenty years. At Yahoo company, also fast search and transfer here in in Trondheim and Norway.
And also working on embeddings, neural search, all of those things. Right? Leading up until ChatGPT, the ChatGPT moment, like November 2022.
And then there was some kind of cookbook, I think, from OpenAI where they said, okay. This is how you can do connect ChatGPT with your data, and here's embeddings. And I think then a lot of developers, right, got into this this is how we can build This is how we can do Rag.
and Rag that it had to be vector embeddings. By the way, I have a small role in that. I I actually was the one who wrote the, Chroma example in in the opening up cookbook.
You did? Okay. I was a was an angel investor in Chroma, before they became a database, and then then, you know, I was just helping out.
Jeff and and Anton from Chroma. I mean, I think Anton left, but I think they've done a great job at promoting retrieval for AI and infrastructure, and they did a did a lot of great things. So I I really enjoy talking to them on on X.
Anyway and then we had the whole vector database. I think PineCoin was one of the pioneers framing it as a new infrastructure category. If you need to work on embeddings, you have to use a vector database.
And naturally then, if you wanna do anything in AI, then you need to have a vector database. And that was my primary motivation for for writing that piece and looking a little bit back, you know, what happened and where we are now and how I see it. And yeah.
So that that was the pure pure motivation.
Okay. And the the general thesis, I guess, if you wanna just sort of recap that, like, you know, I I think it's a very fast rise and fall. Like, Pinecone was a dominant player, that for a long, long time.
And, you know, I don't know my exact sources because there's a lot of rumors going back and forth. But, apparently, they went up to, like, a 100,000,000 AR very, very quickly to raise a big round. And then suddenly, a lot of people started leaving.
Like, suddenly, it's went went from cool to uncool very quickly, and I don't understand why. I don't understand that either.
a little bit going back to their core messaging. If you go to their website now, it looks more developer focused. It's not the memory for AI.
It's not, like, enterprise ish. It's more towards developers now. So I tried to think that they're to go back to their original roots.
I think that's a good thing. But also, of course, there's been a lot of competition in this space. A lot of new companies, one of the upcoming stars is Turbo Puffer, kind of same SaaS model, a little bit different pricing.
And they really talk to developers. And I'm not saying that the companies are dying. Right?
I'm just saying that the separate infrastructure category is dying. Right? Because you have vector search capabilities in almost any DB technology nowadays.
Right? And you have it also in more traditional search engines like Elasticsearch, Solar, Vespa. So I think there's like convergence on on features on both parts.
So and then you have things like pg vector in Postgres. A lot of people, you know, get confused. Okay, I have already a DB, it has vector search, why do I need another DB, like a vector DB, so the whole database concept.
So I think those companies, I mean, are lots of great technology here, don't get me wrong on that, but I'm not saying that the companies are dying, but I'm saying that the category is dying. So that there's there's this distinction. And I think a lot of people, like, oversaw that and, like, came at me and say that, you know, because they had some kind of hate around some of these companies and they said, yeah, you know, go fuck bank or whatnot.
Right? But I actually say that the category is dying, and I'm actually wanna call these new companies that they are, like, search engines. And I wanna go back to the natural I think that's a more natural abstraction for connecting AI with knowledge and all the arguments for doing Rag.
I think the natural concept there is search. And I think one of the insight I have from I I use Windsurf a lot. I love Windsurf.
They are, like, cascade mode. And if you ask it, like, what are the tools you have available? And, like, list, like, seventeen, eighteen tools, like edit files.
But there's also, like, things like search code base, search the verb grep, and these are like search abstractions. Right? And I love that idea where you like just connect the reasoning model with with these tools that are essentially search tools.
And that can help the agent or the LLM to actually formulate the query, know, should I do a grep, or should I do a more of a semantic search, or should I do more a keyword search, or should I just search the web? So I think that's more like the natural abstraction instead of jumping into, you know, vectors and how you represent, that is more of like a detail, like, of how you implement search. Yeah.
Yeah.
that we fixated a lot on, on vector, like, you know, dense embeddings and all that. And I think now we're, we're sort of broadening out. I will also mention that Chroma, I think from the start has always said they're kind of going after information retrieval and not so much, not so much, you know, the narrow sense of rag.
And I think like broadly, this is the consensus that, you know, like, the, the, the category was, was never really going to be lasting for that long. It was just there was just a brief period of time from my one of my favorite early tweets in AI, you know, this post JGBT phase was, I I I summed up all of the fundraising that happened in vector databases for And 2020 it was something like 230,000,000 in all put into all the vector databases. And that was more than the entire life's lifespan fundraising of MongoDB.
Right. So, like, basically, they cannot all win because they they they've already taken more money than supports, you know, one of the de facto winner companies in NoSQL. Yep.
Interesting. Yeah. I I think also on the mon MongoDB, right, they they brought a new category in NoSQL.
Right? And but nowadays, all the other database players have also caught up. Right?
So now even MongoDB has relational SQL. Right? So there's always, this convergence, but MongoDB kind of it sticks.
But I don't think that for Pinecone that was originally leading that movement, it won't, like, stick in the same way. It's too narrow. It's too too narrow.
But I I I would like to say one more thing about embedding. So people like, okay, Joe, but, you know, embeddings is really important. And, also think that embeddings is really important.
Right? Because you can represent more data than ever before, right? Multimodal, whatnot.
Run into a neural network, get an embedding representation, and then you can move this embedding representation around in vector space and adjust to your domain or whatever you're doing. So it's really important, but what happened was that it went mainstream, right? It went from these big tech companies like Google, Yahoo, Facebook, all of them, you know, been working on embeddings for a long time for a lot of different tasks, right?
But with post ChatGPT and we got the embedding APIs from OpenAPI, it suddenly became mainstream. Like every developer would start using embeddings, right, and what to do with them and similarity search and so forth. So so I'm not against embeddings, Embeddings are here to stay, it's just that it's not only about similarity searches in this kind of embedding space.
And then I think more people actually realize that, you know, you actually need something more to it than just embedding and the cosign similarity to do search well, like things like freshness or authority and all of other signals, like, that that really plays into role in in web search. And I I remember one of the OpenAI guys wrote like, you can embed the whole web and then you can build the next generation web search.
just looking at semantic similarity, that's not gonna play out too well. So, yeah. I mean, they're trying to sell you their their model.
Right? So what I'm gonna do is say those, very hype y things. Yeah.
I mean, the the way that I put it is always you're always gonna wanna do a a hybrid query. You always wanna add metadata and, like, you know, do other stuff. I think my question to you is, maybe a very ageless question, which is, should they all be the same system?
Right? Your search system like, Elasticsearch is typically like, you duplicate your whatever your, main, storage of record is, And then and then you have that search index that is basically almost a complete duplicate. Like, you just copy over the documents.
Do you believe in that? Do you think there's a convergence here? This is a fantastic question.
I think for a lot of use cases. Right?
this great extension, p g vector. Right? And I know that I tweeted things about p g vector that was true in the start around the limitations of p g vector.
But there was a rally around p g vector, like adding new algorithm going introducing actually two algorithms, both EBF and, HNSW, adding half vec, adding binary vectors. So actually, what you can see, pjvector is doing more in the capabilities of vector search than some of the real vector database players. Right?
So if you're only looking at, like, vector search capabilities and you already have your data in Postgres and you're operating at a reasonable scale, I think it's fair to use Postgres. Sorry? I think it's fair to use Postgres or use one database if you're, like, operating at you're not operating at a really large scale and you do, you know, some vector search related workloads, and you also use a database for other types of workload, then it might make sense to just keep the data there.
But if you're actually building something that really depends on search quality and your business depends on it, yeah, then definitely I think that you should consider, you know, actually using a real kind of retrieval or search engine to represent the data there. Right?
Yep. Yeah. And is the search system how closely entwined is Rexis and search in your mind?
Yeah. And that's the thing with embeddings. Right?
because embedding based retrieval has been used for a long time in recommender systems, like large scale recommender systems like, you know, TikTok or or Yahoo News or things like that.
Apparently, TikTok, published their REXIS, recently, which is kinda interesting.
Yeah. It's like a cascade of there's always in this system that operates at a really large scale, there's always a cascade of different stages where you first have to retrieve over the candidate pool, and that's typically using, embedding based retrieval. And then you have layers of re ranking layers that finally you end up with 100 candidates or something like that that you actually present to the user.
So I think definitely there's convergence in that embedding based retrieval is also no more common for such systems. So there's convergence there on how how it actually solve on the technology spectrum.
Yeah. Any other thoughts on like, I guess the confusion for a lot of folks who are newer to this, right, who are they understand now that, you cannot just have embeddings only and cosine similarity only. It's just the sequencing of like, what should I do first?
What should I do second? What should I do third? Everyone says like, you know, re ranking is like super important, but that, you know, it it adds like maybe like three to 4% to your results.
And maybe that's not like the the lowest hanging fruit. So I'm always trying to figure out like, what what should I recommend to people? Right.
Like that they that they should start with like you know, like a Postgres or MongoDB as their their transactional and in the in vector store. Then they could split it out to maybe use Elasticsearch or, Vespa. I I I don't know if, that was there'll be the recommendation there.
Redis, I think, is also trying to push themselves there very, very hard. And then you add Dorexis. Is that is that a good sequence?
Well, I think it's it's really hard to come up with general recommendations, like, without knowing what you're doing. But if you're, like, looking to build a Rag application, like, I think most people are interested in in some something related to Rag. Right?
When you have some data and you have to transform your data, I think I think first, I think is it Hamel that always talks about look at your data or, you know, everyone's talking about look at your data. So first of all, you know, how to get your data in in a in a cleaned up way if it's like PDFs or whatnot to do that. There's there's things there.
I think actually that a very strong baseline is the classical b m 25, like, algorithm that's been around for thirty years. Right? It's keyword matching, but it offers a very useful baseline for a lot of different search use cases because it it gives you that baseline.
Right? Then you can start looking at using an off the shelf embedding model to also embed a model and all of the engines, more or less, most of the engines have some kind of hybrid search capabilities, start to play with that. And then if you can afford it, both from a latency perspective and a cost perspective, you can look at adding like a re ranking layer on top of that.
How you stitch that together, depends on, you know, your your framework of choice. But I think most of you you can you can stitch this together, like, with multiple different APIs depending on, your budget, I guess.
Yeah. And, you know, I I I always tend to recommend people to do this offline as much as possible, like batch line, whatever. Most people don't need fully online systems.
Yep. And that that's a friction point because I've been used to working on kind of constrained online systems, like, at at pretty significant scale. And when it's always like everything is online, needs to be a low latency, and then I have problems adjusting to, you know, when you wanna do things at a at a much lower scale.
So I'll give you an example, like calling out to some kind of embedding API to get JSON floats, it's not something you want to do if you're running at thousands of QPS. You don't wanna add that dependency, you wanna have something local, something that is faster. So I've always been like, okay, I'm going to call out to this endpoint, it's going to take three hundred milliseconds to get this large float.
It's it's something that I like, oh, shrug. But now I'm shifting towards more, know, it's easy, it's API based service, you don't have to think about it, it's just there. So it's much easier to build from, right, to have something that is API based.
embrace that mind mind. I see. I see.
No. And so when I say offline, I mean more like, not in the critical path, like batch Yep. Systems because yeah.
And it's interesting.
the models alongside of the database. Are you are you bullish on that kind of stuff? No.
I'm not. I'm sorry. I I I'm not.
I I think this is Yeah. We also seen other players that tries to, you know, move a lot of the logic into the database, agentic embedding inference and whatnot, LLMs. I think it's the right direction is to keep infrastructure a little bit separate from that because they're, like, different scaling properties.
I I think people can stitch those two things together instead of trying to do everything with one single platform. So no, I'm I'm not bullish on that because I don't believe in the developer experience of writing like these huge SQL statements for transforming data from this and then embedding it and then writing it back and expressing this in the database. Like, what does this do to my database?
Is it like calling out? Is it what's going on? I I I I tend to want to have more control over cost and performance and what's going on than just writing some really large SQL, to to execute.
It's interesting. I I think there's this constant tension between what should live in the database versus what is an external system. I don't think it's a clear cut.
Like, you know, classic was, like, like, the the the cron service, which we have in in the super base. Okay. So cool.
Like, any other, like, hot takes or, you know, what are the biggest criticisms that you got after you published this? You know, like, you know, what do you agree with? What do you what do you disagree with?
Yeah. I think one of the things that people pointed out, if something goes semi viral after a few days, you discover that there's a lot of replies that you didn't see, and you're like, okay. But I think one of the things that stood out was that people said that Joe is saying that Rag is dead because of vector database infrastructure is dead.
Right? And I think that was a misunderstanding as well. And I think that comes from people making the connection between Rag and vector databases, like, it's so strong.
So when I'm saying that the vector database infrastructure category is dead, it's like, okay, Rag is dead. And I I think that Rag is definitely not dead. Right?
Augmenting AI with retrieval or search is still gonna be relevant, and I think it's gonna be relevant for a very long time. Right? So that was one of the things.
repeat. Every time. Every time.
I'm just like, you know, for me, you know, I I put out this, like, sort of cryptic tweet. I was like, you know, this is Llama four is gonna reignite the rec long context versus frank debate, but it it will actually resolve the debate, but not in the way that you want, which is Hold on.
is too cryptic to me.
No. It's just like there's there's they're, like, five other guys, like, saying, oh, like, you know, long context skills rag, like, RRP rag. And I'm just like, guys, you you are you're you're idiots or, like, you're you're engagement farming, basically.
Like, a lot most most likely, you believe like, they know what they're doing, and they're just saying nonsense just just to just to have fun. And, like, there there are, like, people who don't take them seriously.
Yeah. No. But but they also, it's nuance to this.
Right? So I've seen people do rag when there's no need to do rag, meaning that, you know, if you have one PDF, like, with visual information and things and you wanna chat with that, definitely that case is probably like that if you don't have, like, high QPS and things like that. So I think there's nuances around this that there's definitely I had a call with someone that had, like, 300 articles, and I said, you know, this will just fit into the context window of one of these Gemini models.
You don't have to have a vector database for this case. And they were so surprised when I said this, oh, can you really do that? Yeah.
But it's also look at it. It's like just we had like four k context window, right, and now 10,000,000, right? So and that's fast, right?
And then people are still running their initial demos from early January two thousand twenty three, right? So where you were dealing with four k or eight k. Right?
So so some parts of it is still not is is not relevant now because we have longer context windows. But I think retrieval, of course, it's gonna be there for a long time. One example I love to bring up is, like, one of these small toy datasets from Trek COVID.
There's, like, 170,000 documents, and it's already 36,000,000 tokens. And you're not gonna load all all of that, you know, for a single query. Yeah.
Awesome. Do you have a take on knowledge graphs in GraphRAG? I think that the GraphRAG well, I have a lot of takes around it.
I think that one issue I mean, graph databases is a database that kind of solves one particular problem, and it does it well to traverse the edges in the graph and, you know, random access and and jump across. But the core issue issue is actually to build the knowledge graph. Right?
The the entities, the relationships. So if you say graph graph, you know, databases or graph rag is gonna kill vector rag and all that discussion, I think the first issue is to actually build the knowledge graph the first place, right? And if you use a search engine or a dedicated GraphDB to actually speed up and accelerate the searches, okay, fine.
But I think people like, okay, if I'm gonna do GraphRAG, then I need a graph database. And I hate that connection between doing something and then connecting it to some specific technology. And I think a lot of people do that.
Right? You jump from some concept into some technology. You can also do graph exploration with a search engine.
Right? You can so this you don't need a specific technology to do it. And can graph rag be better than vector rag?
Yeah. For sure. In some cases, it might make sense or hybrid or whatnot.
But I think people get caught up in, some specific technology all the time.
Yeah. But, like, I I think that's okay. But I'm still trying to validate the presence of knowledge graphs in LLM applications because, obviously, with LLMs, it is much easier, better to create these entity triplets and all that.
So theoretically, it should be better. Yeah. I mean, in the past, knowledge graph has been a dirty word, but now maybe it's not.
yeah. Maybe I think with LLMs, you can do a lot more things around data generation, you know, in general. So generating those triplets is a bottleneck.
Right? It's been a bottleneck. But now we have LLMs, so I agree.
You know, now it could be easier to actually build what matters. Right? Which is those triplets.
Yeah. Okay. Awesome.
Any other opportunities that you find? I I I know that you you mentioned Gina. I think they're a prominent European, start up in in Rag.
And then I think over here, Voyage just got acquired by NVIDIA. Now anything on the embedding side, like, do we need a lot better embedding models? Do we does what we have from the big labs good enough?
Oh, I I hope to see I mean, Voyage was really leading the pack on doing domain specific embedding models, like legal PDFs. And what I want to see is more embedding models in that direction where you essentially, you know, represent this PDF as an embedding or multiple embeddings, you know, for legal domain or finance or health. I I hope to see that grow so that you can have a better starting point than just those text models.
And I've been a huge believer in using visual language model as as a backbone for embedding models, where you essentially take a screenshot of a page. You don't have to go through OCR. So you you you then get a much richer representation.
You don't have to go through these complex processing pipelines. So I hope to see more innovation. I'm not sure if it's gonna happen because I think it's a difficult business model to be in, like Because you have to have an API based service, and you have to do batching, and you have to make up for the compute, and then, you know, are people willing to pay for it?
And and I think maybe that's why Voyage got acquired. I think also GINA is doing a lot of great things in this space now, especially in European languages. But I think every company is trying to move up in the value ladder.
Right? They want to move into enterprise search or move into a different direction. So, yeah.
But I but I do hope that we will see more and better, like, general embedding models.
Yeah. Yeah. I mean, I'm sure I think the Voyage guys are very happy because it seems like they got acquired for a lot.
Yep. Okay. So okay.
Cool. Anything else before we we wrap? Any calls to action?
Any, you know, parting rants on the topics of the day?
No. I would love I I mean, if you wanna connect with me, you know, for the audience, you can find me on Axe. I'm under the handle Joe Bergen there.
So I love our show. Shout out on on Axe. So I hang out there quite sometimes.
Yeah. Yeah. I mean, it's it's where the AI community is.
You know?
I'm always trying to, like, grow on LinkedIn or YouTube. I I mean, there you know, there's a lot more people there. You know?
There's Twitter is always this this, like, echo chamber.
Yeah. But it it's, it's not the same. I mean, X has I mean, we wouldn't have this meeting me and you without X there.
Right? So it's a great place for really high signal to noise. And I think the AI community there is is really great.
Yeah. Awesome. Well, thank you.
Thank you so much for having me, Vix. This has been awesome.
Shared via Hopper