Executive Interview
Building the future of AI: Classic ML to GenAI w/ Patrick Wendell
Patrick Wendell · Co-founder & VP Engineering · LinkedIn
80% of the postgrad uh table.
>> Those cannot be slow. Those need to be very very fast.
>> It's not easy.
>> It's not going to happen like super quickly.
>> Hey guys, today we have the pleasure to host Patrick Condol is a co-founder and vice president of engineering at datab bricks. He also helped launch data bricks in 2013 and he was also one of
the founding commiters of Apache Spark.
Hi Patrick. Hi Quinton.
>> Hey thanks for having me guys. I'm happy to be here.
>> Yeah thanks Patrick. Uh so maybe let's start with the basic. Can you can you tell us a bit more? Well let's say what did you work on when you start working on data bricks? Uh I saw you walked a
bit on on Spark and then what are you doing now and what was kind of the journey to get there?
>> Yeah, good question. So I let's see. So my career history is pretty simple because I was at uh I did my undergraduate at Princeton and then I went to Berkeley to and the reason why I
went to Berkeley was to work with the team that became the founding team of data bricks. So I I was familiar with their work and um and they were doing very interesting work in the space of
kind of distributed computation at the time. At the time the work was more about scheduling and resource management. It was this project called Messos that was really long time ago so
people probably don't know that largely but um but that was interesting work and then they were just starting to get more into the data analytics space. The general problem they were looking at was
dealing with very largecale computation and and that had recently become much more commoditized, more available because of some hardware and software innovations that had happened. So, so I
thought it was a cool team and they were doing interesting work and then I joined that team and then I only stayed two years at Berkeley before I was doing a PhD but then we left to start data
bricks in 2013.
Um, so I didn't get very far into my PhD. We became more interested in the commercial aspects of the software. We actually tried to give away the software and get the existing companies to use
it. So Spark was just starting when I went there and I helped build the initial versions of Spark and and funny enough, we tried to give it away to people. Um, we went to all the company
because we were just doing research. We thought, hey, you know, let's get let's get people to use this stuff so we can write some papers. So we made everything free. It was open source and then uh a
few companies uh would would use it. A lot of the end users would use it but then the vendors that are the ones that control the distribution of the data analytics software didn't want it. They
said oh no we have better stuff. Our our products are better. And uh and we said are you sure it's free? You don't really don't want this. And um and then it was actually for that reason that we went
and started the company because we felt okay well if no one else is going to you know commercialize this stuff then I guess we're the only ones left to do it.
So so I left to start data bricks with that team in 2013.
And uh and since then um you know in the early days we were all just engineers.
There was nothing else to do other than engineering. We had to build the initial product and come up with the product vision and uh try and win over the customers. And then now we're a much
larger company 10,000 person company.
And so my job is to support you know and and manage large teams of engineers. Um so I so my day job today is that there's around a 500ish person team that focuses on our AI and data science uh product
areas. So that's my official job and then supporting that team and then unofficially um as a founder you kind of never lose the job of founder uh ever. So, so that is like a whole other job and that
involves company, you know, product and culture directions and leadership hiring and occasionally crisis management. Um, it's just like an all-consuming uh job.
You never It's kind of like being a parent. Like maybe your kids move out of the house, but you're still kind of their parent always. Um, if that makes sense.
>> Yeah. And I think what's what's super impressive like it's it's like being a parent and it's your first kid and like the kid is doing really really well. So I guess congrats on that.
>> Thank you.
>> Yeah.
>> And I'll be curious. You mentioned that you're working with the uh the co-founder. So how do you split the area between the uh like the co-founders?
>> Yeah, I would say all of us have roles that are a little bit different. We we certainly lean heavily on product and engineering. So that that's really our background and our passion. And in my
view, a company ultimately is the product that it sells. That's really what it is. Sometimes people get confused. They think, oh, the company is like the or charts and the you know, no,
it's really about the identity and the soul of the company is the product. Um so so most of us stay involved in the product and engineering side and as we scaled we brought in really strong
leaders to lead the different functions.
So like um you know we have head of sales, head of marketing, um head of product. I think you you interviewed Adam Conway. I saw uh I saw that interview. So um so so that was also part of the uh important important for
us in in making the company successful is that we realized we have to like bring in other great people and let them cook so to speak and let them be great and and build out these functions and we
have to decide what we as the founding team want to focus on. Um, and we're always happy to jump in and help. Like I jump in and help sales whenever they ask. But, uh, but you know, there's
people that that's their full job and that's what they're doing. And, uh, within the founders, I would say we've each kind of found areas that we're particularly passionate about and and
competent in uh, as you've seen like Reynold is very very focused on our core data warehousing and and, uh, SQL analytics kind of product portfolio.
That was what his PhD was in. That's he knows that space extremely well. and mate and I focus more on the AI side of the house uh day-to-day. Um but it's not static, you know, we we each are just
jumping in where it's needed and uh and where we can help. So I expect those things will change as they have changed over the years.
>> And and so do you work on the like at the public level, you know, helping deciding what to build and stuff like that or is it more like management of teams and you rely on the PM to, you
know, know what to build at the end of the day? Yeah, I I work heavily on the product definition as well and the product strategy. Um, and I think I think the job of the leadership on so so the job of the
leadership and the founding v vision the founding team is to put create the vision and the framework of like where roughly do we want to go like what roughly do do we want the product to
look like and what are what is the strategy? What are the advantages we're exploiting when we choose this direction? But that leaves a lot of room for there's still many degrees of
freedom within a particular strategy.
And then and then we have to work closely with the teams of PMs and it's not just PMs, it's tech leads also. At data bricks the tech leads the tech lead and PMs kind of work closely together on
the you know filling in all the details of what the thing actually looks like.
So it's an important thing to also leave enough room that the people are empowered and they can you know they can uh they can really define exactly what things look like. But I would say that
you know all the founders are pretty hands-on when it comes to product as yousef before we started recording Ysef was telling a story where I I guess I reached out to you to ask about some
customer use case uh like late at night one day and that's very common uh behavior by me and other founders.
>> Patrick you mentioned that with the co-founders you're shaping the future of the product. So how do you decide uh how to staff the engineering team on which project? I know it is something super
complex to do.
>> Yeah, great question. I I wish it were could be simplified to an algorithm, but it really can't. Um, and it's one of the most difficult uh difficult things because when we look ahead, we see this
universe of things we could build and and many of the things have amazing justification for them. Like we could go in so many different directions and there's great there's customers that
would love that. We ourselves think it's cool and innovative. Um, but the universe of things we can build is so large and it's it's infinite and then the number of resources and people and
teams we have is finite. It's definitely bigger than it was in the past, but it nonetheless is finite. So, um, so then it's this problem of like prioritization and what do you do and what do you not
do. Um, and that is is much more of art than a science. I think some frameworks we use are um trying to do uh the most good for the most number of customers.
So, so if there's a feature or something that particular one customer wants, maybe it's really important to them, but it's not generally re reusable. It's not generally something where we can get
leverage out of it where we can build it once and then has benefit to many many people that we try not to do too much of that you know because because every feature you add it adds also complexity
to the product and that complexity is itself a big cost and a drag on everyone. So so we try to build things that have very broad uh utility. Um, another framework is that you want to
you want to balance um continuing to improve your current product area and then entering new areas and in fact this is like some things that come many companies they build one great product
they bring it to market maybe even they go IPO or they become you know but they're not able to really do the next thing they're not able to do the next act and uh why is that is because you
end up so there's so many great ideas to do within your current uh thing you're doing and um that list becomes infinite and your customers are are loud and they want things and you have to like do
what's right for the customer. So I think uh the the other framework we try to use is is make sure we're we're retaining a certain investment portfolio for the future and I think that's one
thing I think data bricks has done relatively well compared to others in the market is that like for instance we we were really a great uh data engineering platform for data engineering and for some of the machine
learning workloads occasionally but then we wanted to get into the data data warehousing space. This is the SQL analytics space and I promise you we had a full list of things on the data
engineering side to keep doing like we could have kept everyone busy doing those things. Um but we wanted to carve out space to enter a new market and to to attack a new type of product. So, I
think it's um this is another consideration is is how do you balance what you know you're good at and and what you you hear from your customers as incremental improvements you can make
and then entering new markets and t tackling new areas. Um, and yeah, it's just it's just very difficult and requires a lot of practice to get good at.
>> And and how so how do you navigate like purely new market? Because you mentioned DBSQL and SQL uh like if you go and build a warehouse nowadays, I mean it's technically it's it's not easy, but like
it's not a new like you kind of know what to build in a way like you know what the customer wants, but like if you take Genai, it's so new like we don't even know what it's going to be in a few
months. H how do you and and it's so far away from I guess your initial PhD and and what data bricks was uh when you started data bricks. So how do you navigate that? How do you know what to
build and what to do? You know where is the market going like it seems really hard like do you talk to I don't know leaders like how do you manage that?
>> Yeah. So it's a good question. So so first of all I wouldn't even say it was trivial to enter the SQL market to be honest. Um although you're right that it's it's a more mature market. So it's
it's maybe more known what what customers want and how you succeed. Um that also means there's lots of entrenchment. There's incumbents in that market already. There's very mature existing solutions and the bar for
moving from what I have to some new solution for a customer is very high.
You better give me something that's much much better or much cheaper or something. So we needed to have a unique angle when we approached that market.
And and the unique angle we we had was that we have this philosophy of like have a data lake type approach. So you we can give you really fast SQL on your existing in place storage where your
data already is. You don't have to load it into new system and and that significantly saves cost and and stuff like that. So so I would say that even that was non-trivial. Now now in the AI
space it's a it's a very different type of market because it's like a it's a nent market. It's a very new market and that has some benefits uh when when it comes to our strategy which is that there's
not a lot of entrenchment because there is no I mean there's such a um there there is no geni vendor that's 15 years old like there it's not defined that doesn't that's the null set um so that's
uh that that means that you customers are much more willing to try to try things and to experiment with new vendors Um the downside is that the the terminal product shape is not defined like the
terminal product shape of a data warehouse is defined. It's it's been pretty much the the core interface people are using has pretty much not changed for 30 years. The interfaces in
the GI space are just being built. So it's actually more like when we built Spark. When we built Spark, the big data interfaces were not mature. In fact, I would say one of the the main things
Spark did was that it defined a very standard interface for doing large scale data processing. And at the time when we built Spark, there was like many different point solutions. There were
point solutions that were very focused on like streaming workloads and then there was point solutions that were very focused on um you know these kind of OLAP or analytical rollup type
computations and there was point solutions that were very focused on sort of slow batch jobs. Um and what Spark did is we we built a decent enough API that you could kind of standardize in
this API and almost like many of the different workloads would work with this API. That's exactly what SQL was like 20 years prior to Spark. Um, but that doesn't really exist yet in the Genai
space. It's actually more in the point solution space. There's like, you know, there's like vector embeddings for doing some type of uh semantic retrieval.
There's like how do you build evaluations and and and measure the quality of your app. There's like how do you handle memory and conversation management? And there's there then there's this foundational token layer of
like I get my tokens from one of these different providers. Um and those providers have different trade-offs. So the the way that people build these AI apps now it's very much similar to how
people build big data apps in like 2009.
It's a lot of glue and stitching different stuff together and um and the capabilities are very powerful. That's why people are bothering to do this. You wouldn't bother to like do all this work
if you weren't getting something very valuable out of it. But the API layer for how you build these apps is like not defined. And that's certainly the the crux of by the way, it might take
several years for that to happen. It's not going to happen like super quickly.
But um but that is really driving our strategy here in this area is like how do we start to define some of these higher level APIs mirroring like the Spark API or SQL where you can build some higher level
APIs and then and then a lot of the complexity that currently people are having to do is taken care of for you underneath those APIs and and exactly the shape and form of those APIs is
going to take some experimentation. It's not going to be like a slam dunk the first time. In fact, Spark's current APIs are actually very different than the the very initial ones that we built,
you know, and um it took some iterations. So, so that's this is a nent market and and the most important thing is to just listen to customers and try to think about where you design the
right abstractions. That's kind of like what our central challenge is.
And there is and there is like I've seen like a huge push on the genai part but what about the classic ML? I mean is it still something that industries are really using like now nowadays like
everything is about geni ji but I've I've seen few announcement or improvement in the classic ML I'm talking on the general space.
>> Yeah. So I think we took a slightly different tack data bricks than some other vendors in this space. Um we've actually accelerated and invested a lot in classic ML in the last few years
and um and by the way certainly Genai is a huge focus area as I just talked about for a long time but um but in my experience uh classic ML isn't going anywhere and in fact for you know
there's just certain things these generative models are really good at giving you what what is really their core function it's to kind of um map a distribution of human language and and
give humanlike answers to questions and and be like a human and and they're really good at like so so for for applications that you'd otherwise have humans for, you can now plug in these
generative models and they do a pretty good job. Um, but there's lots of types of applications where humans were like never very good at good at them to begin with. Like a very very smart single
human wasn't actually that useful to you in this in this type of application. So there's tons of numeric stuff like fraud scoring, you know, risk risk detection, recommendations where it's actually not
clear like the generative models are the right types of models. By the way, it might be that uh deep learning is is relevant. So there may be neural networks and deep learning. Um so I
don't know if you consider that classic or Genai, it's kind of in the middle. I would call it like um you know just codebased deep learning. Uh by the way there's there's huge uh customers we
have doing computer vision and this kind of stuff. Um so so in you know we've continued to invest heavily in like what you might call let's call it code you know code based or model based uh ML and
and we actually see like significant growth in the in in those as well. Um and what's kind of cool is in some ways the Genai boom is giving all these AI teams out there more budget in some
sense. So like they're able to better articulate to their leadership why it's important to invest in this automation.
Um and that that it's it's almost like a rising tide. All the boats are going up and and and the Genai stuff is is growing but also the classic ML is growing too. I do think there's um there
there's another angle to this too which is that I in the same way that Genai is pretty good at writing code. Our findings are also that Genai is like pretty good at building classic ML
models for you as well. So, so we kind of touched on this with AutoML, our first AutoML efforts a few years ago.
That that was not using Genai. I was just using some kind of parameter sweeps and stuff, but but it turns out that uh if you're trying to do a quick regression or forecast or something like
that, uh Genai can go code up that forecast for you. And and so that's another way in which these things like complement each other rather than one replaces the other is that uh you know
more and more we're working on this data science agent and and for for tasks you're doing where you really a classical model is the right tool for the job that doesn't mean there's no
geni in the picture like the genai might be helping you select and and and even writing the code and running that classical model for you. So, so I don't think it's a either or. I think it's
like they just have very different types of strengths and and I don't think uh people will be using Genai for like forecasting anytime soon. They might be using some deep learning though, some
deep learning models for forecasting.
Does that make sense?
>> Yeah, I think I was about to ask about AutoML and see, you know, what was going on here, but I think you answered that question. I think it makes sense like the kind of old way to do AutoML. Um,
and seeing that as being just part of the like of the AI geni stack as some kind of agents agent system that's going to build all the classic ML really makes sense. I think so like just doing the
glue like that.
I think clearly that's going to be the future.
>> Definitely. I I also think a major area of focus for us is and this is across all of data bricks is like many of the things so the people that work with data bricks they're data analysts or data
scientists or um data engineers or more and more often people that work with data bricks are just consumers in some sense. they're just someone else made a dashboard for them and uh or or a report
and that report might involve models or or forecasting and they're just a person that's consuming it and they're they're using data bricks um all of those can be significantly enhanced with genai. So if
you look a lot of our product surfaces you know what are those people doing?
Well a data scientist is creating the model. I just gave you that example. So so as as I said Jenny can help them build that model. um a data engineer is building a data pipeline. We think Jenny
can help them build the pipeline. We've even built some new interfaces for uh iterating on pipelines using uh generative and and and natural language interfaces. Um a consumer probably just
wants to ask questions. They probably don't want to operate at this lower level. They just want to say um you know, hey, maybe someone shared a graph with me. Hey, this graph only shows, you
know, the last two quarters. I actually want to know the last six quarters for this thing. Um, and by the way, how does this change if I add new customer segments to, you know, this report that
I'm looking at? Um, all of those can be significantly improved uh with Genai.
Initially, I I use the self-driving analogy. Right now, it's kind of like the equivalent of like lane assist, like keeps me in the lane and I can and and you know, uh, adaptive cruise control so
I can uh stay the right distance behind.
So, it's assisting me, but it's not doing everything for me yet. But we're actually getting fairly close to, you know, these higher levels of autonomy where it's like, oh, I can actually just
ask a question and and the AI is just going to figure it out the on the the data the data question for me. Um, and so that's another area that I'm very focused on is and I consider that across
all of the data bricks product portfolio. I actually think the the new interfaces and the new product surfaces are going to be very different uh leveraging these geni tools. Yeah and
and your and just for like you mentioned talked about the DS agent. So this this product just amazing. You can just ask to do data exploration doing forecast for example like you will
it will build the the data model for you. We will do data cleansing. It talks about like flow designer where like can help increase the productivity of a data engineer to clean the data do the
transformation just with prompt uh with with with prompting.
So we are like adding the genai to increase productivity everywhere across the uh the platform and that's what leads me to the question like we we launched many many agents. So what's uh
data bricks strategy with agent bricks?
>> Yeah totally. So so agent bricks um is our our product area for people that are themselves making genai applications. So like uh you know just how we are reinventing the our in products with
genai you know our own customers are as well. So they all would love they all love that their products become much more intelligent that they can interact with users and and so forth. Um and I
think I mentioned to you um the real challenge in this early type of market where everyone is building everything with their own glue and 30 different pieces and pulling some open source
stuff and getting it working is that the higher level interfaces aren't there yet. So agent bricks is like our first basically uh uh set of higher level interfaces and and we we focused on a
few different use cases we see people doing with data bricks. So one of them is like uh I would call it enterprise intelligence or corporate intelligence type of thing. So these are like
internal apps that they're building often to um have some corpus of information um and allow for for high quality Q&A uh amongst their internal employees and sometimes external over
this this data set. So this is kind of we call it um uh a knowledgebased Q&A.
So you have some some uh some knowledge base that could be structured or unstructured data and you're trying to build some like really high quality interface to that. Um and then we have a
few other use cases where we've seen but but the the principle in all of these is like what what we think that the our customers should be focused on is pointing us towards the data sets that
are relevant to this AI they want to build and maybe giving us some evaluations. So saying hey this you know these are good answers these are bad answers.
um all of the complexity of dealing with how you tune the search uh um the search parameters and how you uh um deal with guard rails and how you collect the feedback in a good way like that that is
not a thing our customers should have to worry about that should just be done for them automatically by the system they should be worried about the things that are specific to their use case there
which is things like evaluation and did this get the correct answer and guardrails and all of this stuff so so I would say that that from a product philosophy perspective, the idea of
agent bricks is like let our customers move way faster by by letting them focus on the part of the problem that they know that's their business. And all of this like work of tuning and and
experimenting with different models and sometimes we even do model distillation or model model fine-tuning under the hood. That's like PhD rocket science stuff that the average uh enterprise
should not have to deal with. they should just in the in the same way that in SQL you know they just write some query that explains the gist of what they're trying to do and things like how
you optimize the joins and the network uh um performance of moving all the data around to actually you know manifest that result that is absolutely not something they have to worry about. So,
so I would say Agent Bricks is like our our first go at these higher level AI interfaces for for common tasks we see in building Genai apps and um and we're going to be continuing to evolve it.
Right. Right now, um there's there's four or five that we're working pretty closely on, but I expect there to be more over time.
>> And what so I'm curious what's your what's your view on the like on the way user will create product in the future?
like we see like like I tried cursor and I'm using cursor and cloud code and these tools quite a lot and this makes me think do like is do you think it's going to be at some points uh a world
where like to some extent derig does not even have a UI and all we have is just you know some kind of tool we share and then you just go and you use your AI tool to build everything and you would
say hey >> get me a pipeline to ingest the data and then I want some knowledge you know assistant something something to do something and At the end of the day, data bricks is just giving you all these kind of blocks
and resources to do all the job. But like user would just be using super high level LLM and they don't even go to any of these services. I mean not even data bricks like all all of that. What do you
think?
>> Yeah. So um it's excellent question by the way. I don't have the answer for you with certainty. I have some ideas.
If I knew exactly how, you know, all of AI software will evolve, um, I would be like placing lots of bets and and, you know, knowing that they're going to be correct.
But the the industry is evolving in sort of two approaches. One approach is like completely textbased. So this is like you're you're interacting with text. The interface between all these different uh
software agents is text. By by the way, it's almost captured in these two examples you gave, the the two the two types of approaches. So clawed code is like CLI based. It's like this is your
CLI. I speak text and and that's going to be my interface to you as the user.
Um cursor is like more of a of a fully integrated product UX that is like significantly augmented with AI.
So so we even inside of data bricks maybe have a few examples of both these.
So so for instance LakeFlow designer is really like a pointand-click. you're you're kind of uh iteratively um building an artifact. Uh you are at times entering text or or or doing
generative AI type actions, but you you're in the the midst of like a very rich product UX. [snorts] There's other parts of data bricks where we have Genai features where you're just really like
at a prompt. It's like kind of a teletype back to the old days. And some folks think, "Oh, well that's in the future it's just going to be like one textual interface and that's going to
intermediate like all access I have to other products." I'm a little skeptical that that's going to happen because I think these UIs are very rich. By the way, it doesn't mean that they're not
going to be like AI backed, but I I still believe that like the I the reason why these models all use text at the beginning wasn't because everyone thought that was the best way to build
products. It was just that that's the thing the models were really good at is they're just like really good at doing text prediction. That's basically what these initial models were good at. So
I'm so I see it more as like um an incidental starting point but not where we're likely to end up. I think UIs are very powerful for conveying information and for capturing state as someone is
building up you know a pipeline or building up something that they and I don't think it's going to be super possible to like intermediate all access to software through a text interface. I
I think that's going to be not the way things go in the long term. Um but I don't know it to be I don't know 100% to be honest.
>> And and to go back just to to agent bricks for now um it's an easy way for data engineer and data scientists to build agents but what about maybe the long-term vision? Are we going towards
maybe making them uh easy for business users or end users to build agents on data bricks or it's still going to be for let's say technical people?
>> I think I think there is a path for I think the first thing we want to do for let's say non-technical users is really just question answering and that and that's that's less the scope of agent
bricks and more things that we're focused on in other parts of the product. Um I think if you simply just get question by the way our results on question answering are getting pretty
good like like uh I would say right now um you know we have this Genie product in Genie there's a little bit of curation that happens so it's not like I it's analogous to like a self-driving
car but it's it's on a closed course you know it's not out in the city yet someone designed the course and like you know there's it's safe because you're in this safe environment. So, so the analog
of that is in Genie someone curates a space. They say here's the data sets that are relevant. Here's some example queries. But if you're in that closed course, uh you're getting pretty good
answers right now. Like and some of the stuff we have internally is much better than what's in the product today. So like uh you have to a little bit take my word for it, but just wait like you know
a month or so. But it's not an open course. An open course is like I didn't curate any data sets. I just show up and say, you know, and no one gave you example queries or verified queries, and
I just say, hey, give me this answer.
And the AI is good enough to give you the answer. And more importantly, it's good enough to know if it can give you the answer, and it doesn't just make stuff up if it's wrong. Um, we're not
there yet, but I actually think getting there is a huge I, by the way, we're not super far uh at data bricks. Some of our internal research stuff is getting very close to this. So, so I think that um
getting to like for for the less technical user, the first big win is just can we answer data questions for them and uh and that might take some time but I think that's like a huge
unlock if we can do it which is like really high fidelity because because then uh in in your company you could give data bricks at least this question answering part to every single employee
like why not it's just you just ask questions. Um, I think the idea of a lay person building a model or building a data pipeline or building some of these more complex assets is further out and
that's not our immediate focus right now just because there's so much of a win in front of us with question answering. I would say the way we approach these more advanced um things is more augmenting
already somewhat technical users. So like if you look at LakeFlow Designer, the assumption is the person in there is kind of technical enough to really know what this DAG that they're building is
really doing and you know they can even eyeball a SQL expression and be like yeah does this seem like it's doing the right thing. So I would say that the we're more in sort of augmenting mode
for those users right now rather than than automating fully automating.
>> Yeah. Yeah, and I think it makes sense because like it's it's going to take some years before we can remove all the data engineering jobs and and all these you know requirements. So I think
>> it's just much more difficult problem than like answering a an analytical question.
>> But I I still think we're simplifying the journey. Like if we take the example of someone using lakeflow connect to ingest data from Salesforce. Once the data is in the data lake then you can
use designer to quickly transform the data and then go to genie and use this data to ask questions and give them to the business users so they can ask questions immediately or dashboards and
this could take maybe few years ago I don't know three months four months now you can do it really quickly with some basic knowledge >> exactly yeah and so I think so in that story you told there's still a
semi-technical person that's like setting up the core pipelines. Um they're hopefully doing it much faster and much easier than they could do before. And that's a big one. Like for
instance, the claud codes and curses of the world. You know, they're still targeting the technical they're they're ultimately helping you write code and the assumption is that you kind of know
what that code is doing as the person who's um I still think it's not the point where if you just were completely non-technical, you could just click click click and expect these things to
work. But there were a huge improvement in productivity for those people. So like that's still a very very significant um event that those things exist.
>> All right, I wanted to switch gear and talk a little bit a little bit about the Tecton acquisition uh that I saw in the news. Well, maybe let's start to explain uh what's a feature store because I
think Tecton is kind of uh they created even the naming of feature store. Yeah.
What what is that?
>> Yeah. So um this was the the name that team did create the category basically of feature stores um and it's helpful to maybe define what feature means and I I think it's it can be understood in in other
terms too if people don't like the word feature because it's really it's a fairly simple concept. Um it all started when there when um there were certain types of applications that really needed
a lot of precomputed information about their users in order to make a decision.
So so an example is uh I do fraud detection. Um I have a very short amount of time to decide if a particular transaction is fraud and there's a lot of inputs I would love to have available to me when I'm doing
that check the fraud check. I would like to know um you know how many um how many transactions this particular let's say it's a credit card swipe you know how how many transactions has this
particular person done in the last week and how does that compare to their average number of transactions every week um you know what is the uh what is the location what are the top five
locations this person typically does transactions from and how does that compare to the location that they're in right now you know if someone is uh if someone 90% of their transactions are in
San Francisco, but suddenly they're in Batswana. Like, it's reason to believe there might be something a little atypical about this particular transaction. Um, uh, you know, what is
their current account balance? Uh, are they are they, uh, above or approaching their credit limit? How many other transactions have they done in the last 30 seconds uh, or last 30 minutes or
last 24 hours? So it's very expens when you have something that's on a tight budget time budget like that you can't go and run a million queries to compute all those things like you don't have
enough time. So you have to premp compute them. So so these inputs by the way these these inputs are often fed into a machine learning model uh that then computes the probability of fraud.
So so it will have all these things that you've premputed. Those are called features like so it's a fancy name for like a premputed statistic basically.
And it it's funny because they actually started this team started at Uber and at Uber the big problem that they have was ride time prediction. So like they that was just like one of the major use cases
they were focused on. So when you this is a while back but like when you get a Uber they want to predict how how long I think it's both the the wait time like how long till I get this person comes to
me and then also the time it's going to take me to do the whole trip. It actually gives you both. Um so it turns out that that also is a thing where you can't afford to like do a bunch of
computation as the instance someone you know you have you have millions of people requesting rise all the time and you need to get very fast those so so all the inputs to those algorithms need
to be like precomputed and then refreshed constantly they need to be kept up to date in some basically a cache so that cache is called the feature store and those metrics are called features that's basically what
that means um and it turns out that in classic ML especially for these like fraud detection and and personalization type use cases. Um, this pattern is like very very common and it's actually
really hard to do it yourself. You know, you have to comp somewhere you need to define these features. You need to completely keep them refreshed all the time. You need to also make sure that
they're defined in the same way for your like online application that they are for your analytics. So, so you know offline you might go run you might go run a bunch of analysis of um maybe you
want to improve your fraud detection models. So you want to go run a a historical analysis for the last uh year and simulate you know how you would have predicted the fraud and then compare
that to whether some later ground truth you found out that there actually was fraud or there wasn't fraud. You need to go like recmp compute back in time all these features to do your sort of like
simulated analysis in order to improve your model. So that whole domain of uh of computing these metrics real time, defining them, making sure they're consistent between offline and online,
that is what feature stores do. And uh they're sort of layered on top of the lower level uh um um compute and data infrastructure. So it's funny because Tecton was actually a company that was
largely built on top of data bricks.
They you know they they they built these higher level APIs that then underneath and they did target other platforms, but their most popular underlying platform was data bricks. So so why is any of
this interesting? Well, it turns out in the um in the AI geni space, more and more applications are becoming personalized and contextualized.
So, so you know, geni applications are able to be much more rich and and individual and they almost need to be.
So, for for instance, in our own um in our own data science agent, the one I just mentioned to you that's like helping improve uh data science workflows. Well, uh, it turns out that
like knowing the actions you've just done, the last 20 commands you've entered is like extremely helpful in in informing the decision of what like what the next uh generated output should be.
and um and knowing the the data sets that you frequent that you personally frequently access or the last 20 data sets that you've accessed um because these AIs are able to in a similar way
to those fraud detection models, they're able to take like very recent information that's just happened and then use it to help inform like the decision that they're going to make in
terms of giving you a useful future output. So this this problem of like having very fast and online access to to computed features which is really just like some user related metrics is
actually becoming like really important now in a new way for sort of different reasons because of this genai based applications. So so Tecton was like a really strong team. We knew the team
well. the product was already running on top of data bricks and the type of problem that they know how to solve which is this kind of online serving of these very very difficult to compute um
user metrics is like becoming very very relevant now in the genaii world and for all those reasons and we also just knew the team for a long time and thought highly of them. So so that was one of
the main reasons why we did the did the the the acquisition and uh and yeah so that maybe gives you some sense of why uh this is useful to us. Does that make sense?
>> Yeah. And you mentioned that feature store like the I'm talking about tecton um and how it is related to the agentic workloads and during the acquisition of u of neon. We also mentioned that they
were like I think 80% of the posgress tables was uh created by agent. So I believe we're stitching tecton with with neon and also with the recent acquisition of moon cake. So how would
be the interaction between those?
>> Yeah, makes sense. So so these are all complimentary to each other and as I said a tecton is focused on the um the higher level interfaces and then the logic to be computing these features but
they don't actually provide storage themselves. So so they would always run on top of other uh types of storage. And remember there's like the online and the offline. So there's the offline storage
which is like your core data lake. Then there's online and that needs to be a very very fast like you know this example I gave with with predicting the ride time for Uber or for for fraud
scoring or for like pulling the last 10 commands I ran to make sure that I'm giving the right next result for the agent. Uh those cannot be slow those need to be very very fast and so the the
backing storage needs to be storage that can do like relatively fast key value lookup type things. um until re until fairly recently like data bricks actually didn't have such a storage
solution. So so even our own feature store that we had because we also had a feature store it was not as good as the Tecton one. Um even our own feature store we would like tell you oh you have
to like use your maintain your own online thing like you know use Dynamo or use like Reddus or something you know we'll help copy the data over but and our own customers were like hey why
don't you just do this inside of data bricks like why aren't you like why can I just not do my online serving in data bricks um so so part of the the the um Neon acquisition was about uh basically
um by the way it was kind of interesting we were already working on with similar technology on a similar technology stack. Um but it was kind of similar where the team was really good. They
were they had developed a lot of expertise in this kind of um uh Postgressbased online um database management and it just fits in super nicely with the other things we're doing which are also trending in this kind of
online direction that uh we thought that was a really good opportunity. So, so I know it can seem like this um this this you know jumble of like different technologies and different companies and
what the heck is there any coherence to this but um but hopefully actually my hope is that the average customer is interacting with very high level interfaces and how exactly the the feature
computation and the online storage and the offline storage are synchronized their data like that should not be something our customers need to worry about. they they should focus on
defining at a high level their applications and we should be dealing with all of that complexity. So so so hopefully like people see that as a positive that you know we're in a in a
prior world where you might have had to work with four different vendors and glue stuff together like this is all going to become part of one pretty coherent uh application platform
basically for these AI apps.
>> It really I think it's a perfect way to close it.
>> Yeah. Yeah. Well, thank you guys for uh for having me and I've enjoyed as practice. I like looked at I watched several of the other uh podcast episodes and I learned a lot even though uh it's
people I work with every day, you know, just hearing them in this context uh was was very very interesting. So, thanks for the work you guys are doing. Right.
>> Thank you so much for your time.