Fivetran
Fivetran & Databricks Guide
I love the fact that we focus on one thing make data movement as simple as essentially turning on your lights ingesting some data looks super easy AI ml gen AI all right folks uh so today we
have with ky ky is working for five Tran uh and F Tran is specialized in moving the data around I know a lot of customers are using F Tran at their bricks um kayy why don't you talk a
little bit more about you and and what you do at Fan yeah absolutely thank you Quinton uh it's great to here by the way really appreciate you having me on the show um yeah so I've got partner sales
engineering I've got a small Global team that is working with our top technology Partners like data bricks plus Consulting Partners around the world helping to really put together really
Innovative Data Solutions helping use five Trend better uh across that data ecosystem so I have uh personally been around in the data space for for a good while spent uh spent time at Oracle
spent time at in had dup land with Horton works and the five to six years prior to five Tran I was in the data Consulting space with hashmap and NTT data and uh have really enjoyed uh these
these last couple of years over at F Tran it was an easy transition for me because I was already recommending F Tran a lot to my clients and so it it was a natural transition to say hey all
that speed Simplicity operational sustainability and self-service I've been talking about now I'm representing one of those products in the mix that does that so awesome so can you tell us
can you tell us about five Tran what in just a few words what's F Tran doing and why our customer using F Tran maybe a few I don't know a few use cases a few example with a few customers would help
to understand yeah absolutely it's think of FIV Tran very simply as an automated data Movement platform and and one of the things I I talk a lot about is I I love the fact that we focus on one thing
I think the more you get you know kind of uh you move into different areas it can it can dilute your focus and our our CEO's been very clear that we have a very large remit to make data movement
as simple as essentially turning on your lights you know don't let me don't make me think about it don't make me know what infrastructure is behind it don't don't make me configure anything and
we've got over today 8,000 customers that are doing that very thing with FIV Tran across pretty much any source you can imagine databases applications file systems Event Systems uh you know you
think about what's the toughest thing uh to me when you when you think about datab bricks Lakehouse I've got so many incredible data workloads that the datab bricks Lakehouse drives everything from
data warehousing to data lake lake house to uh data applications AI ml gen AI everything in between but I need there's a common theme that I need for each one of those data bricks Lakehouse workloads
and that is high quality trusted usable data that's easily understandable and it's organized and it's really analytics ready and and to me that's what five Trend does exceptionally well again across any data
source so the easy way to think about it is any data source any data bricks Lakehouse data workload all right sounds good and most of the use case would be around would it be more the analytics
use cases or do are you also doing more real time like you know KF real time use cases yeah it's really anything I the way I put it is if you've got a workload that's fit for purpose for the data bricks
Lakehouse chances are five Tran's going to have a really good fit with our sources because we have over 500 today across just about any category that you can imagine so uh a nice thing just like
the just like data bricks we go across industry so you think okay can I can I help with uh personalization or fraud detection in financial services what about customer 360 or product
recommendations in retail or cpg um Predictive Analytics or machine maintenance types of use cases in a manufacturing or industrial setting and again the the industries just keep going
much like data bricks we're not we're not limited by industry in fact you know you think about the types of data that that we moving into the lake house those those pretty much stay the same it's
just how based on your industry you're going to apply it to each one of those use cases and what's that data bricks Lakehouse workload that you're looking to drive do I need to go near real time
is it more um you know incremental over the course of an hour or a day and we can really adjust to all of those five Trends platform does everything from continuous sync all the way through if I
want to schedule it every 24 hours I can do it that way as well and you have a range of deployment options too if I need to if I want pure SAS I've got Pure SAS if I want to deploy my data capture and data
integration and data movement technology behind my firewall on Prim or in a private Cloud before I send it up to data bricks I can do that as well with f Tran self-hosted private deployment I I
feel like you know sometime a lot of people can think that connecting to a new data source and ingesting some data looks super easy like I can open a connection to my SQL I can open a
connection to Salesforce and just D the data what's the challenge here why do I need something like five Trend to ingest my data yeah it it's it's tough I you know I go back to the uh the days of
Hadoop like I said and and long before that and I think you know I always like to think about can I if I'm especially if I'm an Enterprise can I get standardization and predictability in what I'm doing from a
data pipeline standpoint point because it is the first step to getting to those data outcomes and it's it's more than just building when I ask consulting firms you know large gsis and Si I work
with those those types of Partners a lot and say how many Sprints does it take to build a single data pipeline a single data pipeline generally the answer is somewhere in that four to eight Sprint
range so you're talking eight weeks to 16 weeks to design it build it test it deploy it automate it document it maintain it and then once they get really you know honest with me they go
you know what Kelly actually it never ends because something's always changing that that DML or ddl on a database you know A A A change is happening in my Salesforce application my workday
application my service Now application I'm getting files added to um you know an object store like S3 or something like that that I want to move all this data into Data bricks how do I handle
all those changes to me that's one of the biggest things is refactoring for those edge cases and handling that change detection it is exceptionally hard to do at scale and you try doing it
for let's say at an Enterprise a hundred different data sources for for one one one single type of data source do that a 100 times the amount of people that you have to have is just uh incredibly High
and the time devoted to that is incredibly high I would rather my data engineering team and my application team be focused on Downstream building uh really high quality data uh use building
out use cases in The Lakehouse off of those workloads and not worrying about keeping pipelines up and running and you know kind of keep my fingers crossed and hope nothing breaks uh syndrome so we're
going to get you out of that to give me two more weeks two more weeks two more weeks to build the pipeline just one more Sprint I think everybody's been there and uh we can build uh we we
create these Pipelines exceptionally fast we'll see that today as we as we move into the demo yeah and I can relate to that I know I know one project and it was kind of just dumping a lot of post
gray um database into into their air bricks it looked easy at the beginning and I think it took them many many Sprint like as you just said like it was way too long just to inj some data and
if I friend could have done that in you know just when one like when day or so yeah I I always like to think if if we can if when we pair the power of data bricks up with five Trend and let's say
our you know let's say there's a consulting firm or let's just say it's a client that's picked this up you know we have data bricks has done an incredible job working with fivet Tran on things
like partner connect where from the data bricks console I can connect directly to FIV Tran automatically spin up a warehouse that FIV Tran is going to use my users my roles my permissions all
those kind of things there's really nothing to it I'm moving data in a very short period of time into the Lake County so a lot of little things like that I think that can accelerate that
time to outcome speed to impact whatever you want to call it I need to reduce the I need to reduce my data product and service backlog and I think taking care of the front end with data moving in is
really important in order to be able to do that all right so why don't why don't we give it a try H you have you have a demo already that we can uh used to see how to connect to their bricks using the
the partner connects yes I do I I love demo that's one of my favorite things I'm going to it's okay Quinton I will share my entire screen all right so I've got I may refer to a couple of slides as
we go through here I think what I'm going to look at showing today here let me go full mode on this is we'll set up over the course of maybe the next 10 minutes we'll set up an sap Erp connector and a sales force
connector using FIV Tran to the data bricks Lakehouse and then and you you gave me a tip a few weeks ago when we were talking say hey Kelly check out data bricks Lake View dashboards so I
love Lake View dashboards I I am totally sold on that so may try toh attempt a quick Lake View dashboard over some of that data that we we just moved into the data bricks lake house so that's what
we're going to do and and like you you talked about while or I talked about and you ask about like cross industry uh sap Salesforce those are cross industry and your industry use cases outcomes and
those data outcomes that you're delivering are going to be very very specific to those Industries I have three represented up there but that could be insurance that could be media and entertainment that could be energy
you you name the use case all right so what I've got up here couple more tabs I've got uh the five I'm going to hide that so we're we got a clean screen there got the five Tran UI up here I've
also got uh data bricks uh workspace up here as well I'm uh data bricks and FIV Tran both support azure AWS and Google uh happen to be running data breaks in Microsoft Azure here and one thing I'll
point out with five Trend I've got multiple destinations I'm not limited by just having one destination so in this MDS datab bricks H account Hands-On lab account I've got a datab bricks lab
Unity catalog lab Unity catalog serverless lab and you can see how that's configured here if I go in I've got uh def uh determining what uh use see what Unity catalog I'm using where I
want to process the data all those kind of things are really left up to the customer to the end user the other thing you can see here is I've got a lot of existing connectors uh here across those
different destinations and we're going to go with uh sap Erp on Hana today in Salesforce and we're going to set these up and then I may come back a little bit later on Quinton and talk through some
of the um you know some of the unique uh features as well but I want to get these set up for us and get that moving so first thing if I've got multiple destinations configured for this account
five trainers going to say okay Kelly which one of these destinations with data bricks do you want to move data to so I'm going to go with the unity catalog serverless destination is what I
have it named and one thing I like to point out I mentioned 500 plus it's actually significantly over 500 we're at 545 sources out of the box today across databases applications files and Event
Systems I think I'm going to go what do I want to do first today let's go we'll go Salesforce today uh I will set up this Salesforce connector so this is an application connector we'll do SAP Erp second uh
this Salesforce connector I can name my destination schema anything I want if I come out here what I mean by this is if I come out here to data bricks you see that my Unity catalog here and you can
see these um data bricks Destin uh sorry data bricks databases or schemas that are out here we're going to name this one what do we want to call this we will go uh predictable uh Salesforce to data bricks I don't so
it's going to create this schema if it doesn't I don't you're exactly you you don't see if you look at this under this TS catalog demo um catalog I do not have anything that says predictable here so I want to do
something totally different FIV trend is not only going to create the schema create all the uh tables create all the columns for you initially I don't have to do any of that in data bricks first
uh five Tran is also going to manage all the schema drift associated with uh Salesforce as well so um all that stuff is handled for me automatically I go out five trend for application connectors
five Tran is doing an API integration and I always always like to tell folks that you know every single application Salesforce service now workday you name it is going to have a little bit
different um way that they enable you to interact with that API so we we try to take advantage as much as we can of the capabilities that an application like Salesforce gives you give you we do
things like uh we even handle formula fields and and so forth so save and test that you can see I was able to authenticate I'm going to connect to that API and this had this uh says
Hands-On lab account this is five Tran this is we're setting up a production pipeline here so our connection test passed very predictable and standardized in terms of how five Tran is going to
interact with really any application I'm going to get a list of these uh and I'm I'm going against by the way Salesforce developer instance I've got over a thousand tables in this developer
instance I do not want all all of those but I would like to take 11 tables and there's a there's a very specific reason why I want to take 11 I don't want to do all 1100 because we'll probably uh run
out of time as that massive uh data set over the next 10 minutes is moving data over I don't have that kind of time but I'm going to put 11 tables in here because I want to show some
out-of-the-box Transformations uh using a great uh bricks partner DBT as well as a great FIV Trend partner and I'm going to put these 11 tables in here the reason I'm saying specifically 11 is the models that we're
going to build off of these tables are it requires all 11 so we're going to enrich and transform and build things like uh average bookings total bookings uh lead time all that kind of stuff so
give me this is the longest part of this demo is just picking these out individually I got a couple more task user and user rle and that so that's goinging all the the Salesforce API so it means you don't
have to install any agents no no nothing installed it's great question you're calling the Salesforce API and you can see I'm going to block all other uh every other table here and let me show
you what I got I have 11 tables here everything else is blocked and you notice it says review the connector schema below and select columns to block or hash I can even hash at the column
level if I want to for an individual table if I say hey look all right in my Salesforce Source it's okay if I've got I don't know a driver's license number or some sort of pii type column but I
don't I don't need that in my uh data bricks uh Lakehouse data product that I'm building I I don't want to have that there I want to keep that in the source so I can actually hash at the column
level we also have right now Quenton in private preview we have Auto pii detection which is pretty cool so it will flag me and say hey Kelly before you send this data over to the to the
Lakehouse are you sure you want to do that or even after it's moved if it sees some combination I'll say oh wait a minute are you sure you want to do this and I can make make some changes here so
right now I won't hash anything I've got 11 tables let's save and continue the last thing is this is pretty cool so five Tran's going to move this 11 table data set over for me uh with Salesforce
but okay so okay great I've got my initial data set my historical sync what about CDC what about ongoing replication what about when something changes in a new schema a new table new column is
added do you want that Kelly do you want some of it or do you just are you happy with this data set you want to block that and this is this is part of uh your decision making around your CDC and
incremental change capture I can determine that here I'm going to hit continue and and I'm going to say yeah let's go ahead Okay cool so I've got I've got a Salesforce connector I've got
some outof the-box models that five Tran will build yeah go ahead and build that That's All She Wrote off to the races start that initial sync and I want to go back to setup real quick I mentioned incremental I can go
down to one minute syns on Salesforce I can go up to 24 hours and anything in between I had had somebody quit and ask me uh that was a wild hey Kelly I need to go beyond 24 hours well you know 24
hours that's that's a like usually people want us to go faster than 24 hours you know give me down to ad minute I've never really had it that one person said I need I like every week every 48
Hours said well sorry you're good 24 is the best we're GNA be able to do there so this is going to uh okay so we're now that is you have a uh an active pipeline that is you can see here this
predictable Salesforce data bricks pipeline that is moving data over to the unity catalog server list data bricks instance let's add a an sap connector here quickly and then we're going to go check
this out I got to bring up my sap credentials here so give me a moment there we go all right I've got a separate screen on that so sap not the not necessarily the easiest source to
deal with right I mean you've got over a 100,000 and tables it's hidden behind a netw weav application layer you've got all kinds of um incredibly Rich data sitting in sap Erp you've got multiple
applications you got ECC uh that ECC instance could be running uh with Hana could be running with Oracle could be running with SQL Server could be running with db2 or something else or you might
have sap S4 Hana and uh the latest and greatest with sap I need to be able to get to that dat I might also have some sap applications like concur or other SAS applications and F Tran has a range
of application uh or connector options I should say for sap today what I'm going to show you is the sap Erp on Hana connector this connector Quenton uh leverages a remote function call that
that goes through the uh netw Weaver application layer okay so we're not accessing the database Direct we're going through a netw weaver remote function call we've got an sap transport
uh five Tran transport that runs on that sap uh netw Weaver layer and that's this uh five Tran netw Weaver API download that you see if you've got an sap um basis team they are very familiar
with this this is the way that sap communicates with the external world and and one of the preferred ways to do that so you'll notice I've got an application server and I'm going to copy and paste
these things so I don't mess them up I've got an my net Weaver application server host identifier I've got my Cy number client number my user that looks good RFC HF uh hiva and then I'm going to copy my
password in here also can connect uh I can connect with this uh sap Erp on a Hana connector I can go SSH reverse SSH or VPN today I am going to go SSH and that's that's where you see the
value of f TR right good luck trying to do that by yourself you know you you want optionality here too you might have customers that want to connect a particular way or have particular ways
that their Network and security teams connect and you go well I developed this you know SSH thing for a customer but this next customer needs VPN those things are really hard to do on a oneoff
basis so we do try to get just like you saw with all those sap connectors depending on your database and every you got a lot of optionality there so let me look oh you know what I didn't
do is name my destination schema so let's go fast sap Erp to data bricks and I think that all looks pretty good let's save and test I'm going to connect to that okay yeah so I'm going to confirm my SSH
fingerprint that's cool and I should run through the rest of my tests here I love with five TR SAS I I love to see the big green box there that says hey yes all your connection tests have passed that's a great sign
like things are set up on the back end you are ready to go boy you don't deploy any infrastructure with uh five trans SAS at all this is a little bit different than Salesforce now remember
there's there's no thousand tables in sap it's 100,000 tables so the way we did this we said hey search for the the tables that you want to use for your analytics product so in this case I'm
going to grab a couple I could grab anything I want I'm going to grab uh some a couple of tables around material data within sap let's save and continue the Mt and Mara tables we'll move those over to the
Lakehouse and again that's it I'm starting that initial sync I've got uh that rolling I'm going to go back out quickly to my uh Salesforce so let let me you ask you a question like the we we see that with
the UI what how do you do it in real production uh deployment do you set up everything with the UI do you call Api do you use something like theform you to set up the infra how is it working yeah
you you've really got options it's it's actually a really good question let me go out I go out to our docs a lot and if you go out to our docs page you see this API reference we have an open API that
you can see all the different areas that you can use to uh enable additional automation as well we even have a things like a terraform provider um we integrate with things like prefect and
airflow so it I'm going to say Quenton it depends on your use case and it depends on the you know if I've got a thousand SQL Server database connectors I don't probably want to set those up
individually yes I would want to automate that if I'm doing individual connectors uh I've got a couple of sap I've got a work day I've got a service now I've got an S3 or whatever it may be
then those it's you could tell I mean it's very fast to set these up so you have individual connectors that have a lot of the same connector I would say go for more automation if if it's more
oneoff then you know just setting it up just it doesn't take take long at all to do that does that make sense y makes sense see if we got this uh I was going to sync this one more time I want to
pick up uh Transformations here so let's go out go out to what did we I want to see what did we call these we've got let's see last sync okay we've got predictable Salesforce I think the other one I
called Fast the sap was it fast yeah this one right here which is sinking right now fast sap data bricks and we'll come back and look at the metrics here in a minute so predictable is the first one so I'm
going to ref I'm in my catalog Explorer we don't see predictable we don't see fast in these schemas here let me refresh my catalog Explorer there is well there's fast right there fast sap Erp there's
MKT there is oh by the way I love this uh these suggestions on describing your tables I have used this quite a bit since it came out in data bricks I'm going to accept these are awesome I'm a
I'm a huge fan of uh how uh data bricks is leveraging AI to to quickly enable me to describe uh these these tables that I have I'm going to hit accept on that one too because I I know what these are
they're they're they're very very uh very very solid so I've got uh sample date out here and whatever is in that uh sap uh table that Mar and MKT uh MKT table that's what you're going to get
out here it's going to be in you know typical relational format columns and rows and tables and then I can go and do my enrichment Transformations whatever I want with those at that point um these
relatively small data sets obviously for you know purposes of doing this over 10 minutes but here's that uh sap uh predictable Let me refresh this one more time yeah okay so predictable
Salesforce datab break remember we had those 11 tables account contact event then we also kicked off that quick start um that transformation I told you that quick start transformation if I go out to five Trend
DBT packages right here if I come out here I've got qu over 100 DBT open source packages out of the box that I can use that five Tran is development said hey open open source Community customers prospects Partners
feel free to use whatever you want and I've get these out of the box models that are built that give me like on this Salesforce sales snapshot things like average bookings average days open
average days to close and that's what I want to show you the last thing I want to show you on this uh for the for the integration with data bricks is I've got all these models built I didn't actually
have to do any transformation work the models came out of that for me so if I go out to dash boards go ahead so wait wait so was it running on on five Trend the DBT transformation no no no no those
DBT Transformations run in the Lakehouse fantastic question all we're doing is scheduling and orchestrating those that is native DBT core or DBT Cloud running in the data bricks Lakehouse we don't we
don't run any Transformations outside of the Lakehouse zero so you just select the DBT package and run it from it's going to Trigg out the Jing air breaks that's right that's right for the let me
show you this uh quick start data models we don't this is not for every single uh connector so I've got quick start data models for pretty good range of connectors here though these are the
ones that are basically just click one button and voila I've got analytics ready data in The Lakehouse so that was Salesforce that we used here obviously a relational database that's those are
going to be very custom so I don't have you know quick start data models for that but there are if I look out here in let's go back up to the DBT package Hub uh I've got I've got an sap package
out here it's not a quick start but if I have DBT core or DBT Cloud running uh in my environment F Tran can kick off the on once new sap data flows into the Lakehouse FIV Tran would automatically
kick off those those DBT models to run and so you always have the freshest data Downstream um because those are running automatically in The Lakehouse once that new data is persisted in so that's
pretty cool um all right I want to I'm going to go out to dashboards here real quick and this is this I again huge shout out to you for pointing out Lake View to me because I've been using the
other dashboards which are awesome but I like the lake view dashboards let's let's create a lake view dashboard here and the first thing we're going to do I all I've got to do on this so on
the in the other dashboards you know we're going to run some queries and and kind of build uh standard here I can actually select a table now I'm going to select let me select a catalog we're in
that TS catalog demo and then I think I can select a schema here too I'm trying to remember what I wanted to select here what was my sales for uh predictable Salesforce yeah I think
it was that one right there I think we called it predictable so here's all my models out here those are those are staging models I may not have that persisted out let me jump out here let me use I'm going to use one of
these other ones because I don't know if those models have fully generated let me get out this is the same exact uh connector here let's go sales snapshot see what happens see what
happens if I do this for sales snapshot oh yeah there we go so that is that actually is that predictable um remember we called it predictable predictable Salesforce data bricks that that was it and look look
what it did Quinton so these are this is my aggregated Salesforce sales snapshot model that gives me bookings amount Clos this month total bookings amount closes quarter total bookings so I'm going to
take two of these we'll take total bookings and total bookings amount and let's come out here this is the coolest thing I really I really like this let's build a visualization describe a chart
total bookings okay I like that I got I know that's I know that I use this data set a lot I know that's the right number that is the coolest thing ever 3.65 million in total bookings let's do
one more quick one here it looks so easy I mean from the to the transformation go are you kidding me average okay average bookings didn't like that one let's see um let me go I'm let me go back out my data let me make
sure what I've got here let's see what do I have oh I bet if I do this let's see if that if that does it yeah there we go average bookings bang so that's that I know again that is a
correct number that 202 I'm going to accept that and then if I want to come over here pop a title on that thing all right average bookings and then over here let's pop a title on that
one total booking so I don't know I'm not quite sure how long you and I have been talking uh probably you know 10 or 15 minutes or so but but Quenton I'm going to go back out here we went from
no uh connectivity to either sap or Salesforce we Ed five Tran to persist that data into the lake house and we built a very quick okay quick it's simple Lake View dashboard with that
counter uh with that counter visualization we can name our dashboard here this is uh yeah he could be could be anything you could have you know buy and and then and Sh anybody I'm ready to
publish I'm off to the races and I can view my published dashboard there amazing and it's incredible so that's that's some of the power that you get with going I mean you think about this
in terms of just showing uh showing a group within your organization a Le POC one of the one of the biggest things every time is how do I how do I become relevant in my organization and and and
you become less relevant the longer the time is that things take especially to build data products you know two more weeks two more weeks two more weeks you can get these done very very fast with
five Tran and data bricks and obviously I I would probably want to take this I'd want to do more transformation and enrichment in The Lakehouse but man to be be able to show data outcome impact
quickly easily uh to me there's no easier way to do that than this combo of five Training data bricks yeah like the time to the time from inis to Insight is so small like it's it's really amazing it's exciting a
a and I've got incremental change data capture set up automatically if changes happen in any of my sources those changes are persisted over on whatever schedule I set up and then yeah and then
the dashboard we will refresh out of the boook because everything will be incremental what about pricing by the way any impact on the refresh and the pricing is it is it based on uh on
freshness no we we go we're we're continuing to work on how we uh uh handle pricing in the most optimal way today it's something called monthly active rows so it's changes that happen
in that Source one time during the month that's kind of how we price now historical any historical sync is 100% free and 14 days of incrementals are also free we want to you know be able to
get a baseline for what that uh incremental sync is going to be and then what the monthly active rows are going to be what what rows are changing during the month usually it's it depends on the
data source if you have a highly uh High change type data source you know you might be up in that you know 20 or 25% of your total rows if it's not a data source that changes a lot maybe it's in
that you know two to four 5% it just depends but um that's 100% consumption base that's how we do it is is based off those changes on the rows during the during the course of the month so so
wait so if I have like one terabyte of initial data the the initial sync is free and then I 100% 100% free yeah every historical sync is free so all of those two that we just set up you know
we took a small data set but yeah that's that's uh there's no charge for that at all man that's so good so nothing should you know stop anyone from using F TR now I I don't think so I you know I
recommended when I was sitting in the Consulting chair for those five or six years I recommended five Tran 100 125 times and I I just feel like um regardless of the data source now and
the data workload in The Lakehouse in the data brecks Lakehouse I mean' got what you saw there we've got a really really elegant easy seamless way to do that then if you do have maybe it's a
financial maybe your financial services institution your Healthcare institution say okay as much as I would like to do SAS for my data integration I really want to deploy this behind my firewall
we have five Tran selfhosted hvr private deployment for that so I can put all of my data integration infrastructure behind my firewall or private Cloud do that processing there and then send it
up to the data bricks Lakehouse so we can cover pure SAS all the way to pure self-hosted and so again there's that optionality I think is is really important and it becomes critical the
larger the Enterprise is amazing well thank you kitty I think that was awesome now I want to try to use five TR you know as much as I can I hope that was at full and thanks so much for your time
absolutely Quenton thank you so much I really enjoyed it thank you