Loading Databricks updates...
← All interviews

Executive Interview

The future of Data Warehousing with Reynold Xin Databricks Co-founder

Reynold Xin · Co-founder & Chief Architect · LinkedIn

The future of Data Warehousing with Reynold Xin Databricks Co-founder

Transcript

pretty big deal because a lot of the workloads are very spiky can do it line by line we definitely more drinks super awesome because very easy to use we think the next phase of data breaks and

the next phase of also just data and AI Platforms in general is to uh make it dramatically easier for everybody hey folks so today I'm with you SEF and we have a special guest with

us rold chin thank you so much for being here with us today rold hey thanks for the invite and uh glad to be here yeah so rold you are the one of the co-founder of their bricks uh now

working as Chief Architect at a bricks you also kind of co-found Apache spark um so worked a lot on distributed system uh I think you worked a lot on photon the new uh kind of their brick Spark

engine which is rewrited in Victor i++ to to provide faster execution um you also used to be one of the top one Apache Spar contributor um I think it's not the case anymore so which make us

wonder you know what are you working on these days yeah um so many years ago I think maybe in the beginning of daily breaks probably up until 2016 2017 I was writing a lot of code uh on the open

source of spark project I was basically other than sleep E I was writing code um the and and then I kind of stopped cuz I had to spend a lot of my energy elsewhere also I still spent a lot of

time on spark but a lot so on the so writing code um and finally so of couple other people maybe surprised me um last year in terms of numberb commits the um I spent most of my time um now actually

uh basically managing and driving um a whole bunch of our different products in the I would say the foundational data layer um this includes data Brak SQL our engine our storage efforts um also saw

our latest aibi efforts um so a lot of different things um that I'm spending my time on um so it's less so on any individual given projects but on the whole collection of them that's in the

data space okay and Ral so today we want to focus on the dataware housing and can you share a bit uh about day toare housing Journey from how it began with spark jobs and now supporting this low

latency interactive queries yeah um so one thing that's pretty interesting is um we were kind of convinced by our customers that we should be going to data warehousing uh data BR SQL now I think it's about two

and a half years since it's a ga and we disclosed publicly it's uh north of $400 million of AR annual recurring revenue and making it one of the fastest um sub enterprise software um in so maybe

growth that's certainly data breaks history but probably just in general the um so nor of 400 million AR um making potentially one of the fastest growing enterprise software out there certainly

data break history um one of the thing that's really interesting is we didn't actually set down and say hey let's build a data warehouse it was actually about four or five years ago almost at

this point that when we look at uh we we did we're doing a lot of data signs on hey what are customers doing in terms of our workloads and we realized hey there's actually some non-trivial amount of

usage on running bi workloads and in particular also SQL workloads on top of data brakes and this despite us uh making it very very difficult to even discover how to connect bi tools um on

uh to datab BS um and and we made it we off fiscated docs quite a bit because we felt hey data BRS simply weren't great for those workloads in the past um so we didn't want a lot of customer use it and

uh then coming back and say hey how come it doesn't work that well um the but despite that we're seeing substantial uptake in that usage so we started interviewing a lot of them and ask them

um hey why why are you doing this why are you connecting your bi TOS we never really like none none of the solution architect were even pushing for that and all the customer told us a very similar

message which is hey I'm already using datab braks to uh and it was before we pushed for the term lake house so they weren't using the term lake house yet and we're already using datab braks to

generate a lot of our uh data sets they we are doing our injust from data in using data brakes for most of our data sets we're doing the Transformations on them um so the data sets already there

yes we often export them into like a rest Shi or some other cloud data warehouse for the bi work but we are not going to duplicate all of our data sets into a cloud data warehouse from data

brakes and every once in a while somebody needs access to a column or a data set that's simply not available in the data warehouse they're designed for bi um so we give them access to data BRS

that's how it happened um so to us that was a pretty sort of a important um product management opportunity which is hey um there's a very natural fit um for running bi workloads on top of uh data

braks on top of a louse as a matter of fact after we spend time thinking about it we actually think we could do uh do this even better um than cloud data warehouses can on this bi workloads and the SEL workloads

uh we can do it cheaper we can do it faster and we can certainly do it and and the intrinsic value of that having access to your so almost all of your data um in a way that centrally govern

is very appealing um so we sort of double down on it we build a pretty sizable team we s of revamp a lot of things to make it work much better uh like today what we have I think it's a

very compelling offering that's U great performance uh for both very long running quers complex workloads and very uh for very short quers as well um and um so that's May SQL and and what was

the like the what were the main changes which really made a difference you know from like from running jobs to doing low latency like was it Photon was it more like there are so many um obviously

there's real big ones Photon was very key in that uh it was a complete revamp of our engine um and it's MPP benine and then there's also um Unity catalog itself is very important because um when

um when it's just a field data Engineers building very complex pipelines you don't necess need a lot of governance but once it's Enterprise wide what give you out to a lot of teams um analysts

governance become extremely important and uh cus was also a pretty key change uh being able to spin up comp in seconds um and have them ready available um it's a pretty big deal because a lot of the

workloads are fairly spiky um and the uh being able to basically uh satisfy the elastic Compu demand um is pretty important but there's also there's tens of thousands of very small changes that

have gone in even just in the case of latency Photon was very big deal but we had this Laten SE tiger team uh whose job was nothing but just let's try to shap shave off latency as much as

possible regardless of where it is in the stack um and they've done so many little changes um to uh that adds up to a pretty sizable game and every year we see engines becoming simpler and more

performant where are we now are we like reaching a plateau or do we have more tricks to speed up the the computation yeah um we definitely have more the uh the there are so many things we

want to do that we haven't been able to got to um and I think we will have easily the next 5 years um cut out for us potentially much longer but I could very easily see hey here's a road map

for the next 5 years when it comes to engine performance point of view um the I think there's a lot that would go into uh even lower latency I think we could probably squeeze out another 100% so

basically things even 2x faster um there will be a lot of hard work um as you know so of the more um the faster you make it the harder it is to make it even faster it was relatively trivial 10

years ago to be Hado by like 100x it got a little bit harder to be spark Yourself by like 10x and it certainly got harder to be the better spark with Fon and now we have to beat the current version of

photon um but yeah I I think there's a lot one pretty important important thing is um it has to do with Simplicity also is um AI um there are a lot of knobs in s of complex performance engineering and

uh many of the knobs I think maybe not all of them but I think vast majority of them will be able to do away with AI and having machine learning models actually learn based on um your actual workflow

um and make predictions based on hey what maybe what's the best parameter um to set um and the I think that would still also get us a next time so so when we say AI U like are you thinking about

things like the query plan and knowing you know I don't kind of what to join first based on some rules on that thei would have learned things like that yeah certainly the query Optimizer is one

sort of component I think can be disrupted by AI um for those of you that maybe have a little bit more knowledge of the internals of a database or data warehouse the Creer Optimizer is the one

that decides that basic decides how you should actually execute a specific query and historically query optimizers are very uh very dumb um they're actually very complicated but in practice very

dumb it's wrong all the time I think I make this joke two years ago at data an Summit U about hey if uh career Optimizer is often wrong like 10 orders or six orders of magnitude wrong um if

any human being in the uh attendance is wrong six order magnitud times um they'll be fired but if it's sec optimizes is wrong um that we'll have to learn how to deal with it um the I I

think because how wrong uh it can be and how easy it is for it to be wrong um this is actually a great opportunity for machine learning models to come in and actually do better this is a little bit

in state of art research SL um nobody have cracked the code there's no sort of wellestablished literature on here's how do you do it I think we're at bleeding edge and we could potentially be the uh

place to solve this problem because we are one we are one if not the largest um data platform that runs all like we see the largest amount of queries and we have the largest amount of data just by

the nature of um sort of the data engineering side um so U we could actually uh leveraging all those workloads to determine hey how do hand we actually build an AI based career Optimizer rather than the traditional s

of cost based optimizer yeah that's I was about to ask you know because you mentioned you have like at least a five years road map of optimization but is it is it mainly kind of known tricks from you know kind of

the standard database world that we just want to reapply to Warehouse or is it brand new R&D that we really have to redesign yeah I would say it's a combination of both there's certainly

some traditional um tricks that are relatively widely known that we haven't even done yet um of the way we do them is depending on the workload cuz I think many traditional tricks they were

optimizing for a highly optimized data layout that customers have spent a lot of time time doing but in practice uh um we don't actually see that um like if you if you s read a sigma the vldb paper

which is the most popular academic conferences for databases a lot of optimization about hey here's like when when when you have an integer columns very dense um how do you join the integer column

with another maybe sparse integer column uh in practice customers don't encode data in integers and think Super clearly about here's how you model the data they just dump the data in often in strings

in Json in whatever the most horribly inefficient encodings possible um so we didn't spend as much time um applying a lot of the more traditional um techniques they're well known and many of them make would make

TPC pretty fast um and that's why we you look at data braks we actually are probably not the most efficient when running just TPC H benchmarks uh we're much more efficient running customer

workloads but we're not actually the most efficient running tpch because tpch a very old Benchmark that's been exploited um to death by all the databases um but we spent more time thinking hey if it's like if all the

data are strings when even for dates and years uh customer don't use integers how do we make those fast um so there's a lot of that there's certainly a lot of traditional things like when it comes to

hey if I do spend the time to create a super well modeled um schema how can I squeeze the most performance out of it I think there's a lot of things we can do there there's a lot more we're trying to

think out of a box every day on hey given this type of workloads they are not very well document in academic literature because um people don't have access to workloads um what other

techniques we can do to actually make them faster and and what about uh kind of AI functions but for end users I'm thinking about you know Max sum average uh but the same logic but apply to text like

can you can you give me an average of the text of you know the full column and the is going to do some summarize operation behind the is it something we are explor yeah we already ship a whole

bunch of those as a matter of fact that um pretty much any function any model you can call like any model you set up in our foundational model API layer as part of M AI stack you could call it

from datab Brak SQL um so this in you can do anything summarizations the sentiment detection we build in a whole bunch of them but you can also just give the arbitrary prompt of saying yeah yeah

I mean you're right we can do it line by line um but I'm thinking about I've seen a lot of use cases where you have like many rows and you want to do kind of aggregation operation on top of that and

it's It's tricky to do um agregation as in uh like like I don't know based on all the reviews extract the top 10 reviews having negative having you know like a negative feedback for example so

you have to read everything kind of do some agregation and you know locally analyze the text and get you back the final results yeah the well I think one thing that's the beauty of all this physical function

is making it very easy to combine unstructured data analysis with structure analysis because you could actually do on a line by line basis extract sentimental extract how well written a specific thing is and assign a

score and then you can actually sort it just using vanilla SQL or you can run arations based on some sort of uh let's say you assign a sentiment score but one of the other thing we're doing

um and sols are still um cooking sols are already available is want to make it uh a very common um type of work people really want to do is to be able to get prediction Revenue prediction cost

prediction usage prediction all kind of prediction into the future um we are we already have those prediction functions in DBC code you can basically call to plot the time series um but we also adding um

a uh we we're going to make it substantially easier to leverage any IBI U but that's a different topic um and I think that's when it be it will become super powerful when we can really think

about how do we holistically design the entire stack vertically versus just working on hey how do I make the data warehouse itself or the engine itself super great or how do I make a separate

so bi to should yeah and uh it feels recently that the market is reimplementing oldfashioned transactional databases Concepts like we have liquid clustering Auto stats storage optimization stored

procedures are we going to fully catch up with the traditional SQL databases at some points uh with Integrity constraints isolation levels yeah um I think by the one one correction I think the question makes a

lot of sense there's one just one correction liquid doesn't exist anywhere else um liquid clustering is actually one of those things you ask me hey are we just implementing no tricks uh or is

are we inventing new thing liquid is actually brand new invention that I don't believe any database in the past have had it um and clustering everybody has it um but traditional clustering has

a lot of flaws including um typically they have this major minor uh sort of column effect which is is the First Column is specifi have basic the biggest impact and the second column have

substantially less only matter if you uh say go to if all all they have the same First Column um liquid's Beauty and um the Brilliance in it is it doesn't matter um it would actually balance out

regardless of where um so the position the column specified in um and it's also super Dynamic it makes it um you don't have to worry about sort of skill on one uh part and so like for example if have

very popular key liquid solve all those problem and liquid is like not a catchup per say um but the uh so the theme of our question is really about hey will we uh maybe say 5 years from now we have

all the features from traditional data warehouseing databases in the lake house um I think the question is approximately yes um what we want to do is is there's a lot of good ideas in traditional

databases and no reason not to have them um I think transactions a great concept um I do think there are better ways to build data pipelines um with more rigor without transactions basically

leveraging idency but I think transactions a pretty useful concept A lot of people are getting used to we should definitely have it um the uh interior constraints um we already have

it in many places um I think many of those are great Concepts that a lot of workload depend on want to make those workloads very easy to migrate to a lake house so um if we don't have them yet

we'll work on them um we very likely support them there might be minor the reason I said approximate yes there might be some that you know every system every database have some feature they

added and everybody is like why do we even have this in the first place nobody wants to use it it's super difficult to understand I don't think we'll get to those uh but for all the major ones that

people depend on um and are useful we definitely should have them and I see no fundamental reason why we can't have them at some point I think the lake housee should be a super set of any data

warehouses yeah and and so like now we still need to have the kind of real time database I'm thinking about the MySQL post gry you know this kind of tech and then you have to always

copy everything dump everything into lake house to do your analytics do you think at some point we'll come up with some system uh you know being able to do both and you just have the data at one

place and you can do low latency um I'm thinking you a mobile phone application Level and on top of that also like the the the full the awarehouse stuff do you think it's doable yeah

um I think it's doable like but but there's a term for this which htab databases um the um and but if I slice and dice it like I I'm a the opinion that htab databases in a specific form is not

super useful and I explain that um I think there are many if you so the world is largely separated into hey here's like the analytics stack which is slow slowly changing um and uh your runs maybe you

touch a lot of every query touch a lot of data and are somewhat slow like a fast analytical query might be 500 milliseconds um slow one could be days right and then the other side the oltp

databases online transaction processing databases this is like the post my SQL updates are super frequent you can run maybe a million updates a second and slow like ques are run in like 10

milliseconds um and a second query is considered long running um I I think there are some fundamentals s of trade off between when you design the system to go one another um and the so in an absolute sense a

Tren tradeoff is not really possible to design one system is run super well for both but um my my theory with system design is tradeoffs usually exist at a boundary like at the frontier of the

optimization space um that's when you really have to make tradeoffs but vast majority of software system on the planet I should put in efficient um there are not efficiency well important

goal is often not the number one thing you want to optimize for you want to make sure you can maintain the software you want to make sure the software is understandable so there's actually we're

not at the frontier in terms of efficiency um we we're not too far away but we're not right there so there's often things you can do without tradeoffs uh like the classic trade-off

for example in many system is latency ver throughput but in practice you can prove both latency and through at the same time um for vast majority of systems out there but uh the reason I

said I'm not a believer in htap system in a very specific but htap stands for hybrid transac hybrid transaction analytical processing um is I think when people think about agab they often think

about hey I have a database and I start uh my o I start my oltp database and then I directly run analytics against the oltp data the reason I don't think that is a great idea is in practice um most

analytical workloads um you will want a slightly different schema from the oltp uh workload um that's why the process I mean the process of CDC exists like people um do this so change data capture

they actually encode the data slightly differently when they want to do analytics because the type of queries that running are fairly different so if you need to re you need to change the

data anyway um that step um or that duplication of data doesn't seem to be an issue now why do you care about having exactly the same system the hood I think today you do care because those many things are

kind of very difficult uh first of all it's very difficult to get data from a olp database into a data lake or a data warehouse It's a heavy it requires a proper data engineering team in many

cases at scale um I I think that problem need to go away in the future um the other one is hey uh maybe your LP database and your uh uh transaction workloads or your analytical workloads

have very different capabilities so now when you deal with two different systems um you are constantly switching back and forth between hey this is a set of capabilities of my analytical side

this is a set of capability of myp side and something overlap they all run more or less SQL but then everything is different they have different data types they have different of capabilities some

have integrity constraints some don't um and that made it very difficult people to switch from one to another um so I think in the future um we should be able to make that difference dramatically smaller and to

make it so lower the cognitive overhead um yeah I think it REM makes sense the one of the common use case I've seen if when you start creating data apps uh and you want to have a kind of a mix of

everything U but I think the the point you make really makes sense if you like if you keep two systems but make it super easy to synchronize them and maybe with things like Aon and all these you

know leg flow connect setup um then you're right I don't see a reason that to have to like don't forget about the lak house Federation which is sometimes for Dimension like the dimension tables

often are the ones that people don't change um and as a result you can just do real time lookups um those you could actually uh either Federate and I think we'll come up with newer Technologies in the future and

newer product offerings that you don't even have to worry about it basically we should be able to blend the line as much as possible even though underlying architecture might be quite a bit

different yeah yeah for sure you could have everything into Unity catalog and it could be like two system and of the hood all right and what are the challenges that you participate uh regarding balancing

the complexity um while you do the integration of like the latest features to make sure that you will maintain ease of use and uh keep keeps having good performance for end users um

the so we started we don't always do the best job um but we are really trying um hard in internally and often you'll see actually in our internal meling list some like some engineering team come up

with a new way of doing things they say hey here's like a config to uh turn it on and I'll turn it off and we'll be okay with that during a preview period uh because they want to be able to get

feedback but you'll almost inevitably see the push back from uh so some other engineering team or from the founder saying hey we should do no up um and the basically we think the next phase of

data breaks and the next phase of also just data and AI Platforms in general is to uh make it dramatically easier for everybody uh yes it's very important to make things more efficient because all

this workloads are ship large and they consume a lot of computer Cycles um and we continue to be doing that but really to uh truly democratize of data in AI uh we have to make it substantially easier to use um and that

is the key and easier ease of use refers to very holistically EAS of use it's not just hey do you have a dragon drop UI or do you s uh um you have uh this thing in every language but it it really means

hey how when you consider the end to endend um user Journey how much does the user have to think and what are the number of maybe users out there that can benefit from a system without having to

call maybe the most technical users uh for help help um so in a way I would actually rank e abuse higher um than performance um John El to help um s a a very well-known Professor uh who used to

be at Berkeley and went in Stanford uh in the last decade it was famous they quoted for saying uh I think the biggest performance Improvement of them all is when the systems go from not working to

working um so the uh that's what I was referring to if we could get uh the system to be easier to use and now it can benef for more user to those user we're getting the biggest performance

Improvement of them are for them and and do you see so talking about new user do you see new user um still doing SQL as a data analyst are we are we seeing maybe more and more data

analyst doing kind of um analyst and SQ and Warehouse operation using python maybe also using llm to help them you know go faster yeah uh all of them I think um there's so on one hand I think we'll

expand into um so people that don't even know SQL um I think um the true democratization is how do we enable every user out there regardless of technical sophistication so they can get

some value out of data now of course the ones that have absolutely uh no maybe uh s of tech data Foundation is not going to be ingesting a paby of data and then um processing that but there's a lot of

data analysis tasks they want to do every day they're not necess satisfied by Excel or Google spreadsheet um that we should be able to help them in the future uh but uh coming back to the

python side I think one of we talk a lot about dbsl and about SQL here uh and um if you look at if you talk to 15 years ago me um I would tell you hey all data analysis are basically done in SQL

um but having seen through what how much of world changed a complexity of a type of analysis have been all the organizations been doing and also the rigor people want to bring um I think

python is becoming um as important if not even more important than SQL and a lot of it has to do with uh sort of the complexity of the task people are trying to accomplish but also bring in

engineering to data engineering if you think about data processing and yes sqls actually pretty powerful it's a t complete language can do a lot of things you can probably even approximate

pretty much all python in SQL you probably implement a python Tri in SQL um very slow one but the um uh SQL is extremely difficult to test um SQL is uh very difficult to moize if you think about the last 30

years or 40 years uh there's the discipline cop software engineering we don't talk about as much this day is it uh but there are a set of techniques um that you so we have figured out to use

to build robust applications uh with software engineering uh and this are cicd uh abstractions um there's a lot of them idees and all of this um but when it comes to data processing um even though we have the

term data engineering uh if you look at it and when it's using SQL most of the techniques for uh uh serious application like software engineering don't apply like how do you unit test SQL good luck

uh it's very difficult to Define like to modularize SQL yes you can create views uh commentable of Expressions but those are very clunky compared with classes functions in uh so python right so I

think one of yes there's a lot about complexity and being able to handle more complex task but I think one of the so maybe the primary building of python is hey it's a real programming language in

which all the software engineering uh toolkits come for free um and by providing the first class python API basically the py spark API the facto standard for uh data processing at this

point um you can actually do very serious data engineering you can collaborate with a larger number of people you can build robust pipelines you can test it you can very easily do rollouts

um staging de prod um this are things it's extreme it's still possible and there are Frameworks out there like DBT that help you with that but I'm much more difficult much more unnatural with

SQL and uh that's I think one of the true power of python and I have a question on the aibi before we close I think like it was unveiled on on the recently like during the data n Summit so aii had two op we

have let's say AI Genie which is super awesome because very easy to use I mean I think it's one of the easiest product to use in data brakes and we have dashboards which is also I think a

revamped version infused with AI because even me who is very bad when it comes to building dashboards I'm able now to build dashboards and I keep discussing with chiao and Miranda but I want to have

your own Vision about what do you and what do you where do you see uh aii dashboards I don't know for the next couple two years what's your vision for a dashboard side or a in general uh

let's say dashboards yeah I I think we'll continue making the dashboards a lot more uh so dashboards are a reasonably welldefined product um and there's a lot of bi2 out there they're more or less basically

dashboard uh tools um the sometimes have people ask me what's the difference between bi and dashboards like they're very similar but uh maybe when you have enough feature sets you become a bi to

uh when you don't have cross filtering and all those features your dashboard um I think there's a long tale of features will continue being sort of implementing so this would become more and more

compelling offering um where I am very excited about um obviously is the ability to combine it with Genie um and um to infuse more and more of the AI capabilities into this um and I think

that I think we'll have new use cases that like today people think about AI capabilities in BIOS they think mostly As of hey let me I call both on AI uh which is what can use that term at the

data in isino um they basic think about how do I add a co-pilot to my dashboard to help me build dashboards um but I think ultimately um business intelligence is about so of getting insights out of data

and dashboard is a form factor for that um so the part I'm really most excited about is not it's just adding um AI capabilities or co-al like capabilities to a dash but really hey how do we think

from first principles and what are the product service areas that we are building that we have we haven't even imagined yet um that will help people get insights out of data leveraging AI

like I think um if it just take a slightly longer term Horizon um the the rise of the internet or rise of mobile phone enable a lot of use cases people never even thought about before those

that yes everybody thought I would be able to call people anywhere I wanted to that was the first uh s use case that people can think of and I think it's very similar to co-pilots or S assistants for

dashboards or a text to this um but over time a lot of use case showed up um like taking a picture geot tag it upload it to Instagram uh Social Network all this new use case showed up that people like

none none of us had imagined um when mobile phone first came out I think there will be a lot of this and uh we will spend a lot of time talking to customer learning how we can help them

and think about some creative ways of uh building this in the aibi portfolio that makes sense so Reynolds when we last time when we asked you to record the session with us you asked us

to raise the bar and to double down the number of subscribers so now we are at 3K what if we hit 6K so can you help us get another co-founder sure absolutely I wrestle some wrestle

Patrick some into this once you get this person will be the third person since we already interviewed mate now it's your turn so thank you thank you so much reyolds thanks for having me thank you bye