Orchestra
Orchestrate Your Data
hi uh today I have Hugo from Orchestra we gonna talk a bit about him and then about his company nice thanks JF hey how's it going all good all good I'm very happy to meet you yeah likewise likewise um
yeah really been a big fan of dat bricks for a long time um I guess bit of an intro about me so you know I used to be head of data at a fintech company in London um called codap and you know we
were we were using data bricks but it was also at a time we were using stuff like Snowflake and you know we had all of these tools right and it was very difficult to just kind of make sense of
what exactly was going on in all of them but also to orchestrate tasks together you know obviously you can trigger things um in you know data ingestion tools you can monitor streams you can
leverage plat forms like data bricks to you know trigger and operate series of notebooks and SQL but then you get into the Realms of DBT and you know you start having questions around how you can
bring in aspects like bi into all of these workflows um so you know basically decided to found a company uh it's called Orchestra and the idea is it's a piece of software that can turn a
platforms like data breaks into pure endtoend um end to end data and I AI product platforms so giving you visibility and control not just over stuff happening in data bricks but also
data ingestion you know streaming data into object storage um you know alerting business intelligence dashboarding um you know data science that's going on in other bits and you know other pieces of
cloud infrastructure and so on interesting how does it work is it is it a python SDK does it work with some yamal files that you need to configure yeah so it it's quite interesting
because um you know where you've got open- Source workflow orchestration tools it requires you to write an enormous amount of code and maintain an enormous amount of infrastructure um and
you know we had a lot of infrastructure already maintained right so we had app Services data ingestion modules you know python running in ec2 and ECS instances we didn't really see the need to tightly
couple executing python code with orchestration um where I used to work so Orchestra is just orchestration and observability I would describe it as um sort of code first GUI driven so you
primarily would develop in a UI um which is pretty Snappy but under the hood we're just building yl files um and you know in time for our you know some it's an alfha at the moment but for some of
our very early customers we're exposing that yel so if you do need to sort of control stuff by git and bya code you can amazing can you show us an example like how to maybe orchestrate a workflow or I don't know
something on databas also another tool as well yeah yeah yeah absolutely um I will I will share my screen just go YOLO and share my entire screen that's an Inception I see my my my my face twice yeah I know sorry about
that um but yeah I don't know I I feel like when I was a when when I was working in data the thing that I was missing was like having having a sort of person watching my back right it's like
as a as a as a data engineer you're always sort of working for other teams but who's kind of watching your your back um this is sort of the goal with Orchestra so in terms of managing things
like connections uh you know it's a very very easy and straightforward way you take credentials from your under platforms ensuring they have as little permission as possible um and you know
connection is very straightforward right give it a name get your access token get your workspace crack on um you know we support multi-tenancy So if you wanted to have a different sort of account or
you know for different people right maybe your engineering team separate your analytics team um you know you can do it that way but the beauty is really in stitching things together and
building pipelines with all of your component bits of infrastructure quickly so if I go ahead and look at our little data breaks module that we've got running here we can see that we're doing
some data engineering right uh we're leveraging a your data Factory to copy some data into blob storage we're triggering a load of workbooks uh in data bricks to you know ingest dat using autoloader um you know
bit of couple of sort of quality life quality of life things we do in this notebook here but we're also combining this with some s so you know here I'm using some f Tran to move data into
Delta uh and I'm also using a piece of software called portable for some connections five TR doesn't support um and you know you can see how easy it would be to sort of add something else here right so if you're
leveraging something like DBT or maybe there's a datax one time run it's just a question of you know sort of parameterizing the task and cracking on um we also have the option to kind of
collect additional data from platforms like data bricks um even if you're running stuff in DBT which is pretty cool um and you know in this example we've got uh you know fairly simple
pipeline right it's not kind of bringing anyone into the loop but in Orchestra we give you full observability and lineage of what's going on as well as running things um which is pretty pretty new so
here this is a run which completed 25 minutes ago um you know we have all of the data we have the task parameters the Run parameters the duration you can go ahead and debug it in data bricks and we
can also see that this is this is leading in to more workflows so in this case you know we're actually triggering a high touch reverse yeah ETL sync and also a powerbi data set refresh um for one team and then on the
other hand we've got a load of other a load of other interesting operations going on here um and in this case you know we can see we're trying to do some more DBT and then some more powerbi but
actually this task has failed um yeah it's it's very interesting so I think about like as you mentioned like you can build some specific workflows on data braks and then of course everything that comes
around like bringing the data like in that data ingestion or maybe like once the job is done how to send triggers I don't know can BEP or any other tool this can be handled over here and it's
it's very simple simple way I've like this the UI is very stretch forward so everyone can create the the workflow quickly yeah and I think something we've seen with people using data breaks in a
multi-team context is that it's great for getting everything executed but when you have a situation like this where one team is powering something that another team does they have to write a lot of code to
ensure that there's you know full alerting and communication between them um and that's something which we've really really tried to build out so if you go ahead and check out a pipeline um
you can actually add an alert um that can basically alert on anything anywhere to anyone you kind of call it like AAS software like ass but that's probably bad idea um but you know like you can you can you
can set alert on the pipeline or the task level right and if there's a dashboard that you know someone's looking at you know they care about you can say okay I'm going to send them you
know I'm going to send them a notification if it succeeds or if it gets skipped right and if it gets skipped it means that something Downstream is is kind of affecting it and then you know you can send it to the
Rel Place yeah exactly um and this is powerful right because and just one question can you can you uh um edit let's say the temp for example if you send this do you have any specific template that you can edit
or just sent like alerts like hey this pipeline did fail yeah yeah so um the alerting is quite uh quite quite granular um I can just get I can get some examples up very quickly so I'm just going to do a little
screenshot because otherwise uh I'll show you exactly what's going on in my slack which I probably shouldn't do um but if I share screen again so you can see here um you know we can send alerts to any any channel um
and you know you see that it's at the pipeline level or if it's at the sort of task level get the state a helpful message uh you know this is something which will pass from the error and then
obviously we have links to um the underlying resources in Orchestra so you know it's quite powerful because you can tell people that stuff is in the process of being fixed and update them when it gets
fixed instead of them just going straight into a dashboard noticing something's funny you know sending a few messages and then having the engineering team say oh no no no it is fine there's
nothing wrong with it right um If you're sort of leveraging an endtoend platform like Orchestra to manage your data stack your ly establishing all those lines of communication between data teams but
also teams that are less technical um and you know that that that builds that builds trust which is obviously pretty invaluable I love it it's it's very uh it seems to be easy to use
straightforward so yeah I can't wait to try it and since I'm talking about this let's suppose I'm a new customer uh uh how can I get started quickly do you have any Tri period or how yeah yeah
um yeah I mean something so free tier um you know we get the the data spaces well funded and there's a lot of you know you can get a lot for free in the market um so yeah with orust we have a free tier
you can run um you know a few pipelines at uh you know like an hourly schedule um and you know there's a there's a task limit but it's it's pretty generous I would say if you're a data team of you know one or two people
and you only need a few schedules this is perfect for you that's perfect and I have two more other questions what if someone want to have more information uh do you have any contacts or
email yeah so um I'm pretty active on Lin uh you can find me there hugu HG luu uh or my email is Hugo get orchestra. I'll make sure to add those details in the description of the video
and yeah I I mean I also I also have um it's crazy I started blogging on data just over a year ago um and that's going that's going quite well yeah I've seen I've seen your Medi I keep following
your your PO and your your articles yeah it crazy because like I got a shout out from Ali yeah I remember yeah I remember that was big yeah that was big it was crazy like traffic to us
I was like whoosh um unfortunately we didn't have a data BX integration back then um but like it's it's ironic cuz that was an article I probably wrote in about 25 minutes and it's like the
second one I ever did so now you've raised the bar so high yeah I know the only way is the only way is down and last question for you what's your favorite uh data bricks features feature oh my favorite data Bri feature
um that is a great question um gosh there's so many just one the the one that had let's say most impact or or based on on your knowledge yeah for me the thing that really like blew my
mind um is its ability to handle streaming end to end spark structured streaming for ingestion basic transformation you can have a data science model get trained in real time and you can also now you know serve that
model in real time using data breaks and then do analytics on the predictions in real time that's pretty cool I don't know anywhere else you can do that unless you build it yourself and and I
think we're adding more features because now you can of have the the lineage you have within with unity catalog so for example you're streaming from a source then you're I don't know uh updating
some specific table so this table is being used to train in a m model can have the lineage and the lineage will go through the till the uh model serving part yeah and and then and then the
thing I love the most that you can also use something also called inference Sables where you can have a Delta table where whenever for example you are using the endpoint you can ask question like I
don't know uh maybe house price prediction you're going to enter all the details those details going to be stored in this inference tables and also the the output so and at the end you can for
example track the quality of the model if you have a baseline to compare to and you can also have this integration with Lakehouse monitoring where you can track the data quality the model quality
so yeah that's I I think that's pretty pretty impressive yeah it's super cool um do we have time for a quick question about the lineage yeah so how exactly does it work if I say okay I want to see a graph of
what depends on what at a given point in time like what is going on under the hood there is data breakes basically saying okay for every asset like a notebook or a machine machine learning
model fetch me the most recent sort of queries or operations and then you know calculating it from that or is it driven by passing you know a git repository that's been checked in at that point in
time so we have something else we have so typically you're talking about two things so we have first something called uh the inside tab so the inside tab because we we do have Unity catalog we
can capture some some uh some metadata so if you for example select a specific table you're going to see a bunch of uh information like how often this table is being uh queried uh the most uh the
users who quer this table the most the table that are joined the most with uh this one and ahead we can also find for example where this table has been used in which notebook uh which and this is
everything like related to the audit to the some logs that are generated and data are passsing them and push pushing them okay and for the lineage part that's another thing so lineage part is
is computed uh when the uh when spark is is being run then we we get those metrics and those met those metrics are put where that's where we can find for example that this table has been
computed from from this one and this column realiz on on this one and this is really this is really uh let's say an advantage thanks to Unity catalog because those metadata are handled very well and we can expose them
to the customer so to see in order to simplify how uh simplify their life for example and the most I think the use case that I found most interesting is people that I mean now people sometimes
update a specific table let's suppose they did remove a column or they they did some change in the The Source system how they can trigger or see the uh the impact of this change and then you can
find quickly yes I I need to update this specific notebook it will impact this specific table this specific column this specific dashboard and it goes from from The Source till the till the serving
part can be like an ml model or maybe L VI dashboard and you can find this easily and this is also part of the uh data Discovery and asset Discovery because this is also something something
important cool interesting yeah the reason the reason I asked about it over time is that um at the data intelligence day earlier this week in London they demoed that but there was a
filter for the lineage which was you know it was a time filter so it was like show me lineage for the last half an hour for example um and I thought that was quite interesting as it wasn't that
you know that was not intuitive to me I thought was cool I think I think this one is yeah I remember I've seen this in the lineage part we need to select 30 days and 60 days 90 days I haven't uh I
still didn't like played with it to to see what's behind but it's main we'll have to go away and do some more research yeah yeah yeah exactly cool good stuff thank you thank you so much uh go uh really enjoyed
discovering uh Orchestra and uh looking forward to seeing you again yeah likewise um and yeah if anyone if anyone's looking forward to working out how they can bring data bricks and all
the other bits of the data stack under one place that's Orchestra come find me perfect thank you go cool cheers mate bye