Executive Interview
The Last Data Engineering Video You'll Ever Need with Bilal Aslam
Bilal Aslam · Sr Director Product Management · LinkedIn
you are a data engineer, you have to be careful about Gil because it's going to kill your job. When I tell them about what we're doing, they're like, gosh, this is super exciting. I want to be
more effective and I I don't want to wake up at 2:00 in the morning. You go from being like a light cruiser. You're very agile, very small, but you know, a single shell can take you out to now
you're on the aircraft carrier and just like sticks around you. You know, our goal is that you shouldn't have to be a data engineer to do data engineering.
All right, let's get started. Um, so today we have Bil with us. Bil, thank you so much for taking the time to have this discussion.
Um, so Bil, you are senior director doing product management at data bricks.
I feel like you have been around at data bricks for like a very very long time.
Um I think you worked on many products but mostly like all the ingestion workload work um the transformation the workflow and like all the orchestration I think I remember you on all the work
you did for the for the workflow the orchestration product at the air bricks um so yeah thanks for being here to have you so maybe let's start talking about product management at the bricks um so
well basically what's a senior director of product management like what do you do? How many people are working with you? How does it work? Okay. Excellent.
Yeah. So, um you know, product management at data bricks is um pretty straightforward. I actually think it's one of the most fun jobs we have here at this company with uh apologies to the uh
you know uh to Quentyn and Yousef. Um so product management here is pretty straightforward and what we do is really I think of it as three things. The first thing we do is we connect deeply with
the customer and we are responsible for understanding the customer needs both the obvious needs which customers will sometimes tell us and then the hidden needs the latent needs right so they
they tell us sometimes we have to really double click on what they really really want uh so that's you know the first part of our job um and that's really involves spending a lot of time with
customers talking to them having dinner with them debugging uh you know uh 8 to 10 hours a week with customers is excellent for a product manager. Then what we do is we partner with engineers.
So product managers don't write code, we don't own code. So we have to partner with engineering and uh ultimately we are working with engineering to give them clarity and a vision for what to
build. So we take these customer requirements which can be very abstract, right? Then we work with engineering to make them real. And then finally what product management does is we um we make
the product successful. And this is the vaguest and I think the funnest part of our jobs and that could be anything. Uh that could be writing documentation. It could be enabling our sales teams. Uh it
could be uh talking to uh a particular customer and getting them excited about the product. So really it's all the other things. These are the empty seats in the room. We fill them. Um and to the
second part of your question, I'm lucky.
I work with an organization of about 30 product managers or so. Um and we have an interesting mix of product managers at data bricks. We have product managers who do both classic inbound product
management and a couple of other different types of roles as well.
And I'm I'm curious to know how do you balance I feel like well I'm going to take me as example so that I don't you know I'm not going to say any any customer but sometime I feel like when
when I've been asked okay how can we improve the product I have some kind of a a short-term vision and I feel like most of customers because they use it daily they just they just think about
you know how to improve the product locally and I suspect that's a lot of the ask you have but yet when I see the road map you and the PM have I feel like you somehow always managed to be, you
know, like way more two years ahead in advance. Yeah. Two years ahead. And and I'm wondering how do you balance that um with what the customer is asking which seems like more, you know, the kind of
daily issue they have like a super long term right now. I I completely understand. Yeah. So actually, you know what's funny is um since I've been here, I've been lucky to be here for a few
years. I've actually seen Data Bricks get a lot better at this. So when data bricks was a smaller company and our customer base was smaller we were in fact a bit more let's call it
uh more short term right we were learning and if you if you kind of think about the uh a company has a sort of gradient descents into product market fit so it's not immediate it sort of descends into
that so it kind of optimizes and I think what happens is that as an enterprise software company we have become a lot more mature we have a lot more customers than we did 5 years ago. We have a lot
more use cases. Uh so what that really means is that we've we have gotten better at understanding as a product team and as an engineering team, hey, is this thing going to solve just one use
case or is it going to solve, you know, $10 million worth of use cases? So with that said, I will say that even though we're pretty good at it, we pay special attention. There always a
set of customers who are in fact, you know, let's say we're two years ahead or 3 years ahead, they're 5 years ahead. So we do listen specially to them and the way we do it is we have customer forums
where you know let's say for example uh startups startups are often different from say large enterprise companies and we pay special attention to them because they're doing things that are at the at
the boundary of what's possible. So I think the answer is yes broadly speaking we do what you describe but we also try to make sure we're listening to the next wave.
And as you can just one like as you mentioned you've been here for quite some time and you mentioned that you have seen data bricks changing but is there any other things you have noticed
like or things has changed for the last couple three years let's say yeah know um so actually it's interesting what we haven't changed in maybe I'll cover that super quickly because um I feel like we don't talk
about that enough uh so what we haven't changed in is customer obsession uh we haven't changed in ultimately what our mission is which has always been the same which is uh democratize data and
AI. Now with that said I think what we've gotten better are we've gotten better at understanding the interactions between capabilities. So for example let's take unity catalog right unity catalog is an
incredibly powerful capability for our customers and as it happens it touches everything in the platform. If you go all the way down to the Spark engine, it touches that it touches storage and it
touches, you know, um the user interface, right? So, I think what I'm pleased about is we've gotten better at understanding these interactions and these are really it's a very large
number of interactions between all these different components. I'd say we've gotten much better at it than we we used to be. Um we used to sometimes discover these at later stages. I think now we
have a better machinery for understanding kind of how does this affect that and would you say like it's slowing down the the new features and the releases of data bricks?
Um you know actually the secret is sometimes customers tell us we release too much we have too or not too much maybe that's the wrong phrase but what they really typically say is you know
gosh we have trouble keeping up with your innovation. Now, they say this in a very happy voice, but when you ask them, what they're really saying is, "Hey, give me a bit more control over this."
And uh to give you an example of something we did recently was we added this uh preview portal, and it's a way for our customers to sign up for previews, turn them on and off, and for
account teams to have an interaction with the customers. So, customers can say, you know what, I don't want this capability right now. So, um to answer your question, we're actually incredibly
fast at innovation. I I do think it's it's almost like a C. There are short frequency waves. There are lots of these smaller capabilities that come out. But to be honest, capabilities like Unity
come out every one or two years still.
So these large bets uh like Unity and Mosaic and LakeFlow, they happen every couple of years, but within those years, we have I think we have thousands if not tens of thousands of features that get
released at still sort of almost continuously.
Yeah. Yeah. It's like that's crazy. I feel like it's so many many small teams and everybody's shipping like all year long and then new things and even for us at the within the air sometime I'm like
it's hard to keep up with every everything. It really is. It really is and I think it it it gives the intro to what we we wanted to talk about. I think you worked on data bricks workflows and
in the past it used to be called data bricks jobs if I'm not wrong. And can you walk us a bit on how this product changed? And just for example for the for the folks who are watching this
recording like last year I think datab bricks shipped more than 70 features just for this specific product. So can you share with us like how does it work and the also the some metrics maybe?
Yeah absolutely. Um so yeah so you know the origin story so to speak is pretty straightforward. Uh one of the things we do as product managers as I mentioned is uh we talked to a lot of customers and
the story here is that uh it was soon after I had started at data bricks uh I'm based out of Amsterdam and um you know the first feature I was asked to work on was something called jobs and um
I noticed pretty quickly by talking to customers and looking at data that hey this was a really popular feature that makes sense right if you're using data bricks uh for data engineering for
machine learning uh The notebooks are awesome, but then back then we only had notebooks, but then you want to run things in production. So jobs was the way to run something in production. That
made a lot of sense. Um, and what customers told us was that at that time a job was simply a way of running a notebook or a script on a schedule. You would say, "Hey, run this at 2 in the
morning." And it would run it. It would spin up a cluster. It would run it. It did that really, really well. So it was a very simple feature and, you know, it worked kind of flawlessly. Um, and this
however was not enough. So customers were coming in saying, "Hey, I can start by running this, but then I want to build more complicated workflows. How do I do that?" So this feature really came
out of those conversations. We had about 70 plus conversations and uh we you know we talked to customers who were using airflow and data factory and all these different capabilities and workflows was
really a an attempt at making something that is powerful yet simple. So one of the things we wanted to avoid was complexity. We wanted to make it accessible to everyone. Um and that's
exactly what happened. So um the statistics that I can share are uh this is an extremely popular feature in data bricks. Uh you know we we we're now doing hundreds of millions of runs per
month. So just the scale at which this feature runs is you could probably call it the world's most popular orchestrator just by straight numbers. Um it runs I think at this point 14 different types
of workloads. So it runs SQL queries. It updates dashboards. It runs machine learning models. It does AI training. It does uh you know DBT projects. It does notebooks. It does everything under the
sun. Um and indeed you know and that we shipped more than 70 features on it last year. And uh this I can tell you this that we're going to ship a lot more this year. Um it it's scale the scale at
which the feature works is remarkable. I can't share specifics but um it spins up a lot of compute and it manages a lot of compute uh for all these runs. Now the coolest thing by the
way I will throw out there is that we've added something called serverless workflows and by the way workflows is the f is we changed the name of the capability because jobs was too too much
about running one thing so we changed the name to reflect it capabilities.
It's called workflows. Now, um serverless workflows is super cool and uh the idea behind serverless workflows is that uh in the past you would spin up a cluster in your account and this could
take minutes right now you have to manage subnets and IPs and you have to manage capacity. With serverless you just say here's my workload and data bricks runs it for you. It's kind of
magical. Um just makes everything simpler to manage, everything simpler to run. And that's just one example of something we built. uh and G8 just recently. Yeah. And just one thing like
make sure to watch we're gonna have an upcoming series on workflows with Prashant. Like we're gonna walk through all the amazing features including serverless.
Yeah. Know like the serverless is really amazing. It's one of my favorites. the the the reason is like data bricks is able behind the scene to do you know many many smart things with AI and ML
just to optimize everything and and make sure the customers can run like the fastest or at the best cost and everything like it's it's going to be super nice going to be super simple for
everybody like yeah yeah all the good stuff yeah I must say sorry go no I just wanted to talk about like you talked about a bit about lakeflow and you're famous because you did I mean you're
already famous you did an awesome talk at dice and I think this is one of the coolest features coming to data bricks called lakeflow I think it has three parts uh injection transformation and
orchestration can you talk or talk talk a bit about lakeflow connect and how it may simplify the customer journey yeah absolutely yeah look the idea Idea behind LakeFlow is pretty straightforward, right? So
everything starts with data. Everything you want to do in data bricks or any data platform, you need data. You want to build a cool machine learning model, you need data. Uh you need to do
analytics, you need data, you want to do machine learning, everything. It's just data, data, data, right? Um and the idea behind Lake Flow is that we want to give you a single place and a unified
experience for doing data engineering.
And that really breaks down into three things, right? That's the injection. How do I get all data into data bricks uh in scalable uh price performant way right then how do I transform it which means
clean filter aggregate prepare it and then finally how do I build these production pipelines and how do I monitor everything so you know nobody likes to wake up at 2 in the morning um
so I can start with lakeflow connect so this has these three journeys or these three capabilities mapped to three different products the first one for ingestion is called lakeflow connect and
Um this is super exciting. It's actually a brand new product and uh LakeFlow Connect is a fully managed uh zero knobs injection capability. It has two big capabil two sets of connectors. So uh it connects to
enterprise applications and it also connects to uh databases. So the idea here is that if you have a SQL server, you have Postgress, you have Oracle, you have MySQL today, you're sort of forced
to either buy a separate solution and you have to go to it and purchasing and all that stuff and scalability. This is actually really hard to do because you're capturing change data, you're
capturing it at speed, there's a lot of variability, you have to set up this whole team to manage and run this thing and not to mention the price of this thing, right? Um just to get that data
in data bricks. that what's really exciting is and I I've seen this again and again guys is that a tremendous amount of business data which can unlock hundreds if not thousands of use cases
is locked up in these databases right and these are database if you could just get this data out you could do you know whichever vertical you're in just unlock hundreds of use cases um so that's one
part of this then the other part of this is there are these ERPs CRM HRM right so Salesforce workday Netswuite and a bunch of these other systems they have similar data. It's not transactional data but
sometimes it's data about your customers or about orders or these are the this is the business domain data right um and lakeflow connect will make it super easy to get that in and then finally I would
say yeah perhaps the third part is unstructured data sort of the third repository of data that you want in data bricks is sharepoint google drive ftp and this is you know unstructured data
and this not only actually it's a misconception that this only powers ML use cases There's a lot of analytics use cases and application use cases that can be powered with this data. Sort of how
do you get this data into data bricks and uh the magic trick behind this is really it's not the managed part because that's a convenience right the magic trick here is how do you do this in an
efficient and incremental way all the way from source all the way to destination and that's actually what lakeflow connect does by working with lakeflow pipelines and lakeflow jobs.
Yeah, I think I think it's one of the you know most amazing features we have like when you especially when you're a new customer and you start with data bricks you want to give it a try and the
first thing you do is like okay now I need to have some data to you know do something but oh no no it's on Salesforce how do I it's on I don't know it's on my SQL how do I get the data in
and then it goes into crazy export dump you have to set up DBZM and some you know madness around that yeah so like no need for DBZM no need for like an IT consultant to come and do it you can
just set it all up itself. Yeah. And so c can we maybe deep dive just a little bit more on the three capab capability you mentioned the so let's say I have a I don't know SQL database um maybe um I
believe we have pos coming soon yes so I have posgrade it's running on the cloud what do I set up do I need to install something on my posgr instance how does it work no so uh it's actually pretty
simple um so what we're working on right now is that if you're let's take SQL server for a second right so if you have SQL server and your data bricks your data bricks data plane can reach that
postgress instance and that's actually quite common among our customers because for example in Azure they might set up private link or VPC peering right or in AWS for example as long as data bricks
data plane can reach that database and it's actually quite straightforward to set up with data bricks with our security controls and then you don't need to install anything on your SQL
server Now what you'll need to do is uh because it's SQL server it has certain you need to run a few commands to enable change tracking on certain tables. If you go to your DBA they'll say
absolutely they know exactly what you're talking about. So the database server has to be told like hey either on this schema or on this table set up change tracking and what that does is it takes
you know several minutes uh you know you just run the commands and once that's done your database starts producing a change data feed and what we do is that's it that's all you have to do then
you go into data bricks into lakeflow connect and you just enter the credentials you test your connection you choose what do you want what objects uh what tables or schemas do you want to
bring over and you say, "Hey, you know, I need to put them in this catalog." So, it's a wizard. It's pointandclick. No coding required. Um, and then you say, "Okay, go." And what we'll do at that
point is we'll we'll set up all the infrastructure needed to both efficiently extract the change feed and also apply it into data bricks to apply that into data bricks in an incremental
fashion. Uh so you know if uh uh setting up the database 10 minutes point andclick five minutes and you're done and uh what's cool is the scales this scales from small installations or small
SQL servers all you know we're running some serious workloads with our customers so scales to very large SQL servers as well and uh as you said Quenton uh we're working on Postgress
and MySQL will be coming as well. Uh and then there's like a long line of these uh databases, you know, Oracle and DB2 that are on the road map.
Yeah, it's amazing. And and so the application would kind of the same, right? We do incremental ingestion and then you can create like you would end up with a DT pipeline at the end, I
believe. And then you just say, okay, run that every I don't know every hour.
So first of all, like the data just shows up in Unity catalog. So these are just tables in Unity catalog. You can read them, you can run a SQL query on them, you can join them, whatever you
want to do. So you don't need to do anything else with them. If all you want is the raw data, you've got the raw data. It just shows up magically and it's kind of cool. It just shows up and
it kind of keeps running right now.
Often that's not enough. Often what needs to happen next is you might say, "Hey, my, you know, my Salesforce customer records, there's some data cleanup I need to do. Uh, you know, I
need to maybe get rid of some nulls. I need to do some, you know, um, normalize, uh, cities and names and things like that, whatever that might be. Um and that's where lakeflow pipelines comes in. So lakeflow
pipelines is a declarative way to build these data pipelines. And by declarative what I really mean is instead of worrying about hey uh you know where do I create tables? How do I orchestrate
the data flow between these tables you basically write queries. You write a SQL query or a Python data frame operation and you connect them by referring from one to the other. So you select you
create one then you select from one to the other and then data bricks figures out the rest. So to take this example of you know let's say you have customers and orders coming in from Salesforce you
want to do some cleanup uh they're landing in unity catalog has two tables what you'll do is you'll go into lakeflow pipeline uh lakeflow pipelines and you'll just write a couple of
queries and you'll say make it so and that make it so is actually magically incremental. So as your Salesforce data grows, you know, let's say you're joining these two tables, these joins
will be incremental. Uh and whether it's a pabyte of data or whether it's a you know, 100 kilobytes of data, you don't have to think about it. So we're making efficient use of your uh you know of uh
of your infrastructure and ultimately cost you less. Um, so yeah, pipelines is the way to do transformations and I think it's part of the vision you you have mentioned like and and to be honest
it's it's it did blow blow my mind like when I've seen like okay we have we will switch DT to streaming tables and materialized views and we're going to include enzyme to have this incremental
injection now those features are part of the um are part of the lakeflow connect with the acquisition of Archon or Arion I always struggled I think you corrected last corrected me last time it's Aron
right yes so with this integration of Aron to build the new technology I mean when you start thinking about it say okay I mean those guys are really visionary to have this like thinking to
have three years ahead of us you know and yeah by the way I will say what's what I'm really excited about is that you know ultimately the the the thing that excites me most about lakeflow
connect lakeflow pipelines and jobs is that our goal is to get this in front of more users. You know, our goal is that you shouldn't have to be a data engineer to do data engineering. You shouldn't
have to understand uh you know structured streaming and incremental you know manual merging and you shouldn't have to understand debzium. So in a way we're kind of going from the high
priests of data engineering and we think everybody should be able to do this now.
They should be able to do this in a way that is easy to use. It's safe so you you can't make big mistakes with it.
It's also cost- effective. So these are sort of the three promises of LakeFlow as a whole. Um and I think we have a you know we're I'm excited because I think we have a real opportunity to bring this
to them to everybody who uses data bricks. Yeah. And it's like it's interesting to see how data bricks is evolving like we I feel like we started as like the super kind of advanced tech
companies with spark and now I know it's like bringing data engineering and and data science to everybody. Indeed. And if you and what's cool and again here's what's really cool is if you if you
think about it as a technology stack it's not like we're getting rid of structured streaming. Structured streaming is there except now we're building higher level abstractions and I
think that's really ultimately what's what's happening in the data engineering space is that we're we're building on top of you know structured streaming. So for example lakeflow connect is fully
built on top of delta life tables. So uh it uses all the capabilities we have. It uses apply changes into it uses structured streaming. It uses all the tech that we built over the last three
to four years. It just has an abstraction and in this case the abstraction is hey it's a wizard and in the wizard you connect to a source and it just brings in the data. It's still
the same tech. And what's particularly cool about it is that you get all the governance that you're used to. You get the cost management that you're used to.
We're not building a disconnected product just to get to market faster and help our customers. We're building on the stack they already have. Yeah. So, so I feel like if you are a data
engineer, you have to be careful about Bilal because it's going to kill your job. I hope not. Actually, my hope is that I make you a lot more productive.
And by the way, just so you know, uh when I talk to data engineers, they actually they have so much work to do.
They're being asked to do more with less. And believe me, no data engineer wakes up and says, "Let me write a lot of boilerplate. let me write a testing framework. Let me write a data quality
framework. They have so much work to do.
So when I when I tell them about what we're doing, they're like, gosh, this is super exciting. I want to be more effective and I I don't want to wake up at 2 in the morning. If you if you think
about it, guys, it's very similar to software engineers, right? Like software engineers are like ve well there are few who write assembly. Many of them I write TypeScript like I enjoy running in high
abstractions. It's really that same movement.
Yeah. And I think on top of that with all the like the AI assistance like everything gets you know easier and easier like within the Arabics. So I feel like now the assistance is getting
really good. I must say like the first day it got released I was like it's not really you know 100% working and I tried again like really tried for real project a few weeks ago and I was like wow
that's really working well now and and same for big thanks to Western. Yeah, thanks.
And I start I'm starting to do everything with AI. Like I'm using also, you know, cursor AI. Like it's so smart.
Also, I feel like it's really making data engineering and data science so much easier. And and by the way, that's, you know, just so you know, that's sort of the I I imagine that as LakeFlow gets
out the door and matures as a product, I'm excited about adding AI. I think uh AI in the service of removing toil is wonderful. I think uh less less manual heavy lifting uh undifferiated heavy
lifting for data engineers. Yeah. And all the you know um so right now I think when you have an issue in workflow like the is going to be hey I think I know what's going on you know do you want to
check that and like it's so helpful.
It's like okay yeah I'm going to do that. You need you just need to open the page and go to the assistant and ask question how can I fix this and it will go and trigger like some magic to to
suggest some some code uh resolving or I don't know because you forgot a comma or I don't know something silly and you know bel data bricks I think since 2017 like when you used to only
have uh spark I think back in that time we don't we didn't even have jobs. I think it was you needed to orchestrate this with the notebook call in another notebook. That's right. That's called
notebook workflows. It's still there by the way if it's for backwards compatibility and and now like I was at data and AI world tour and now we we are we're able to build a complete like end
to end application just using data bricks like for example you can bring I don't know whatever data from Salesforce using lakeflow then you can ingest some PDFs or whatever in volume and then put
those package them into train a model and then use a model model serving with the data bricks applica apps Yeah, data bricks apps. Exactly. To have your own hosted app there and without
needing anything from from the cloud provider, just databased technologies everywhere. That's pretty cool stuff.
And I think um absolutely and I think that's generally again these are like higher level abstractions, right? Apps uses infrastructure we've been building for years. LakeFlow uses infrastructure
we've been building for years. So it's exciting. I think a lot of things are coming together. Yeah. And in the past I think the data bricks I think title was making big data simple. I think I think
now it should now it's more making like people's life simple. I think that that's really like and that's like the the abstract one is democratizing access democratizing data to everyone. Yeah. No
I I agree. And I think you know ultimately uh I think all of our customers are you know they're they're they have very busy jobs very busy lives and our ultimate goal is to help them
get that sort of inefficiencies and twile out of the way and can we talk a bit about about I mean I remember when it was announced it was a massive uh was something very exciting for all data bricks folks
because it was the missing piece like how to ingest data from the from other sources.
So how did this integration goes like and how does it work and is it like are we revamping all the code or just leveraging their technologies what's what's happening there? Yeah. Yeah. So
um a little bit about the industry right so um Archon was actually a data bricks partner and they were doing well and uh they were we were seeing customers being very successful with them and um they
are in an industry sort of colloquially known as change data capture and it's a bit of a specialized industry because ultimately what they're doing or companies in this space are doing is
they're in the business of figuring out let's call them obscure database log formats, right? So, databases are not new. They've been around for 40, 50 years. So, something like that, right?
Pretty much every database has supported some notion of a log um CDC log, binary log. There are different formats and they're vendor specific. There is no open source standard. And so, the really
cool thing about Archon is I guess two things, right? So, one thing is they really cracked the ability to read these logs efficiently. So they have a lot of tech uh and now we have a
lot of tech on how to read these logs efficiently. That's number one. Um number two is they really lived this space. So they you know um databases are uh if I think about technology the
databases are on the order of complexity of operating systems. They're very complex pieces of technology and Archon has success in deploying uh Archon and reading from them. So which is
remarkable uh for for a small company.
uh so they've really lived the domain so they know how to do this with Oracle and SQL server and my SQL and Postgress it's a phenomenal team so that's the third thing now what we've been busy doing for
the last year is actually integrating them into data bricks uh so and really that's about making sure that when we give this to our customers that we can operate at the scale of data bricks we
can operate within the governance boundaries of unity catalog and then we can operate in a fully incremental manner from start to finish. Um, what this really means is, uh, sort of
absorbing them and integrating them into our platform, something we never had to do when they were a partner. Uh, but you know, when it comes to their actual let's say the Archon connectors that
actually read data from databases, those are fundamentally unchanged because they've they're battle tested, they're in production, they work. What really we've been working on is making sure
that they are fully because ultimately we don't want to build a separate product. you know, LakeFlow Connect is called a product, but it's really just part of data bricks. Anybody will be
able to use it. Um, so there's a little bit of that trade-off that we do in product management where we say, hey, do we go to market really quickly or do we make sure that we build a single
product? And we're optimizing for building a single product here. I think Dris made, you know, couple of acquisition like the big one, including things like Mosaic and I feel like all
of them u worked really well in term of product and integration. It seems to me like I always had the feeling it's super hard um because like you have company A, they have, you know, their own stack,
their own people, their own culture, and then you have to merge them with company B. Everybody has to be happy, the tech has to be happy. Like what's the what's the secret? How do you make it work?
Yeah. Oh, I think it's always hard. I think um I I give a lot of credit to the acquired company actually in this case and actually in every one of these cases I give a lot of credit to the acquiry
because it's sort of like moving I I had a startup once and we got acquired and it's jarring. It's like there's a spirit of joy followed by absolute terror at what just happened because you go from
you know uh I'll use a naval analogy.
You go from being like a light cruiser.
you're very agile, very small, but you know, like a single shell can take you out to now you're on the aircraft carrier and there's like ships around you. Uh it's jarring. You're going from
something very small to very big. And this this order of magnitude change is to the to the psyche, the human psyche is very jarring. So I give a lot of credit to the acquire in this case. Uh I
think they're mature leaders. I think they're seasoned. uh and they're they've they they anticipated this change and they worked with you know let's say someone like me who's been here for 5
years and I'm used to being on the aircraft carrier um I think also ultimately though it's about strategy right these are strategically aligned acquisitions uh all that means is that
there's a ton of value for our customers when we get this right so we all have aligned incentives and these are founding teams that are excited because we can increase the scale dramatically
increase the scale of their products, right? And ultimately, that's what founders want. Founders want to build successful businesses, successful companies and successful products. And
to them, they just get this massive shot at helping customers, you know, all across, you know, from SMBs all the way to enterprises, the world's fortune 50.
And that's a pretty rare opportunity.
So, uh, I think it it goes both ways.
Uh, Bel, what about the timeline for this LakeFlow product?
Yeah. Uh, the great question. So, uh, Layflow Connect is already in preview and the way we're rolling it out is that we have a small private preview. It's actually not that small anymore, uh, of,
uh, several connectors, uh, Salesforce, uh, Netswuite, Workday, uh, Google Analytics, um, I might be missing a couple and SQL Server. So, these are all available. So, if you're interested in
using these connectors, you know, reach out to your account team um, and, you know, we'd be happy to see if you qualify.
um LakeFlow pipelines and LakeFlow jobs which by the way lakeflow pipelines is the future name of Delta Life tables.
It's 100% backwards compatible. So you you know keep using Delta Life tables everything will work and database workflows uh will become data uh lakeflow jobs. We're kind of going back
to the jobs name because it actually makes sense now. Okay. So these are in active development. We are looking to start a preview uh early next year. So end of this year, early next year, uh we
actually have some customers already using these products. So it's actually pretty exciting. It's a very small group. Uh it's an active development. Um and uh we are actually rolling some of
these changes out. Sort of the backend changes uh without a lot of fanfare. So for example, performance improvements, error messages, and things like that.
We're just rolling those out. We're not holding those back. But look for towards the end of this year pro maybe early next year for for a private preview of uh lakeflow pipelines and lakeflow jobs.
So for lakeflow jobs are going to be the same thing backward compatibility which means if you have your workflow running compatible Mhm. fully backwards compatible and uh fully integrated with
pipelines and fully integrated with connect. So if you have a if you're building something in database workflows today it'll just show up. It'll just continue to work. You won't lose any
history. You won't need to do any migration. It'll just work. Yeah. All right. Cool. So, start start using DT now. Start using workflow. Well, they should be using workflow because like
it's such an amazing product and you will be good. And just for your for your information like we had a talk like few months ago or few weeks ago I guess with armburst. So DT we knew that it was
little hard to debug have some issues.
those issues are being fixed like now we have unified DT pipelines with those like two concepts streaming tables and materialized views with uh enzyme to make sure that you have this incremental
processing we have now the possibility also to publish to multiple cataloges and schemas we are getting rid of the keyword live so and also simplifying the debug experience in the notebook itself
so the experience is heavily simplified exact Exactly right. Yeah, the experience is massively simplified and I think you know ultimately what's um what we learned from Delta Life tables was
that it solves a very hard and really valuable problem and what we're just laser focused on making it very simple.
We're laser focused on making it available to everyone and again this goes to the hey you shouldn't have to be data engineer you shouldn't have to understand spark to debug a delta life
tables pipeline. So, in LakeFlow pipeline, we're giving you an amazing authoring experience, an editor for this with built-in debugging, built-in error handling, assistance, and all that. It
should be as simple as notebook really.
So, uh I'm really excited about that.
Yeah. And we're betting a lot on DT because like like V vector search is backed by D back MVs more than us by the way. More than us, uh thousands of customers are using it very successfully. It's a very successful
product and it's just usage is ramping up really quickly. So now is actually a really good time for us to work on this.
Yeah, I think uh I think it's best way to to give it a try. Test it. If you had like tried it two years ago the experience has been completely revamped. So just give it a try. Trust
me, you're going to be amazed. And a lot of things have been heavily simplified like uh this like sdly change changing dimension. Uh a lot of lot of things have been like done in two lines of code
instead of having to write complex uh structures just to handle some tiny tiny stuff. Thanks a lot Bill. That was that was amazing discussion. uh like I'm so glad we have an improved full data
ingestion data in uh like like flow connect like flow transformation likeation everything within one product it's going to be so helpful for so many customers thanks for your time uh and
well we hope we will have you another time soon I'd love to thank you Yes.