Anomalo
Future of Data Quality
one hi I'm youf and today I have special guest so I I let you introduce yourself sure uh Hey Yousef my name is Zach I'm a solution engineer at anomo and my background is in data engineering amazing very short concise
so can you tell me more about anamal uh yeah sure so analo is a data quality platform uh so it's comprehensive we use AI machine learning to automatically detect anomalies in the data and also automatically find the
root cost so it's meant to make that process a lot faster and a lot more scalable across an entire Lake housee that's amazing so I believe just put some constraints and then uh you
have a job that run behind the scene to make sure that you have the data quality captured in a special way yeah yeah exactly I I could actually walk you through some slides on what that would
look like amazing go ahead all right cool so I I'll just share this um and this will be really quick but for where anomos sits first and foremost is within the data brecks ecosystem we can have lots of
data issues coming from lots of different places so all the way at the sources for what's coming into the lake we could have lots of different data quality issues come up we could also
have those arise within the ETL area so data bricks workflows for example and all the way at the end of the pipeline uh so those would be you coming from Tableau dashboards uh or data science models uh
so where anomalous is in the stack is we want to catch all the anomalies at The Lakehouse level uh so we we filter everything that's coming from upstream and prevent them from propagating
Downstream and so anoma would sit right here now what does data quality monitoring look like uh so to your question youf there are a few different ways in which a Namo is going to be
monitoring data and so in a combination of all these ways we can get a comprehensive approach uh so first you've probably heard the term data observability uh so this is going to be
one very small part of an overall data quality monitoring strategy so this is looking at pipelines making sure data is coming in on time it's fresh the the number of records are complete perhaps
lineage but we actually Le on Unity catalog as well uh but then we go into the actual data quality where we look at the data those would be validation rules which are things like making sure every
record follows a rule if I add column A to column B it should always equal column C uh and so this is getting really granular so we can catch those issues but where this misses is it's
really hard to maintain at scale we can't have a rule for every issue and every column in every table uh then we're we're going to have thousands hundreds of thousands even that's going
to be difficult to manage and create and lots of false positive notifications uh so then next there's a way to look higher at aggregate levels so these would be metrics where we're tracking
spikes and things like sums or averages but we use time series models and so this is the introduction of some machine learning to make sure we're only reporting on really significant issues
like massive spikes are are significant drops uh and then finally this one's a paradigm shift this one's really new that we don't really see anywhere else and this is unsupervised modeling and
what unsupervised ml is doing is it's casting a wide net there's no configuration needed and I'll show an example of that in a bit just looking at the data to see if there was significant
drift compared to historical data so no configuration needed uh this is going to be finding spikes and missing values drops and segments of the data as well as value distribution shifts um and so
the one downside of this approach is is not targeted we're looking for significant shifts and columns across the table using sampling so all four of these types of categories are what make
anomala comprehensive and these are going to be sitting on every table within your Lake housee uh so one last part about a data quality monitor strategy is detecting and alerting are a significant part of
it but the last part is resolving it uh so when I was a data engineer half my time was dedicated to resolving issues so anomal is going to speed up that process by giving a root cause analysis
automatically so we don't have to tell anomalo to look at anything specific if it catches an issue uh like in this particular case it caught uh that if we multiply two columns together it equals
another column uh but that wasn't the case for about 230 records uh well it gener generated root cause analysis and saying okay this is actually coming from New York records and I could go into
that a little bit so the last thing I'll mention on this is that anomalous for all data users so similar to bi dashboards now like back in the day we had uh you know those nightmare
dashboards like Bob J Crystal Reports uh and those are all managed by Enterprise infrastructure um so it was really difficult for people to get new visualizations Along come new newer
tools like Tableau powerbi looker and all of a sudden that puts power into the hands of users or all data users be able to create their own visual analytics and their own reports so Namo is tackling
this in the same way but of course in a governed way so Enterprise can control around this but this could be a Engineers data Engineers data platform uh data scientists anyone working
machine learning uh as well as the consumers all the way down soon stream uh we we see analysts as well as product owners um using anoma to either view results or create their own checks so
that's a a basic overview and I I have a short demo prepared that I could hop into but any anything else uh any questions at a high level that you have no but to be honest I'm amazed
increasing productivity simple to use I mean that's amazing uh yeah yeah no I mean it's pretty cool I I wish I I had this tool when I was a data engineer uh you know especially when I was in finance uh this
could have saved me a lot of headaches and from getting in trouble a few times um so I I guess I could just go directly into a demo I want to start in data bricks just to show you how anomalo
could integrate so in unity catalog there's going to be information on our different tables and you know lineage anomal is actually going to surface the data quality issues it found directly
into Unity so it's saying okay anomalous monitoring this table and I can see that there are few different categories of checks uh most categories pass but key metrics and validation rules fail um so
as a user that might concern me so I could hop directly into anomo to see those results and anomo will display all those different checks and the ones that failed here now before I go into these checks at a high level I
want to take a look together at how we would connect to an Al so this is the first page I'd be seeing if I log into an Alo and this is the tables page what I'll be getting here is a list of all my
tables and Views in data brecks or whatever connected data source uh so if this is the first time someone logs in of course they wouldn't be seeing any tables or views here uh you know if it's
the first time anyone in a company logged in so what I would want to do is connect to a data source and I care about data bricks here so from here I just give analo data brick service
account credentials and whatever permissions that that account has it will automatically pull on a list of all those tables and views right here um so as soon as it does we just select the
ones we want to monitor basic configuration and anom is going to automatically start monitoring those tables so uh what does monitoring data look like uh now again they're going to
be those two overarch categories observability which is a really small part of an overall monitoring strategy uh just in pipelines for example and then there's that deeper data quality
before we hop into those I just want to take a look at these visualizations down here uh what an anomo is giving me here is a profile my data so I get a list of all my columns and Within These columns
a distribution of the records so for example this is credit card transaction data and uh if I'm a credit card company I could see as a user in my Merchant Cate category column that is pretty
common for customer transactions to be related to online retail like Amazon as well as food um such as groceries or restaurants so across thousands of tables I won't be an expert in all them
this will help me get oriented around so let's go into the different checks what does monitoring my data look like uh you know briefly I'll touch on this this is looking to make sure uh the metadata is
representing that the table is there columns haven't been dropped uh the table's being updated consistently and then we got that deeper data quality section which we're making sure that uh
the actual data itself is arriving on time not just that the table is updated and the data is complete so what anama is doing here and again this is completely automatic is it's going to
generate a Time series model to make sure that the number of Records that's coming in on any given day is what it should be expected to be U but a few months ago I could see there was an
anomaly one thing you're probably noticing is there's a degree of seasonality in this data so on weekends uh we can see that there's significant drops uh I mean not significant that's
normal Behavior Uh fewer transactions so anomal is going to note that seasonality it's going to use things like day of the month time of the year even us holidays uh and it's going to adjust those models
to make sure it's not generating false positive alerts uh to catch the significant ones and suppress noise if I'm an end user and I'm getting notifications which anomal can send through slack teams P Duty uh some other
channels uh I don't want to be getting a lot of noise otherwise that's when I check out and at that point it's like I don't have a data quality to it all so uh that unsupervised machine learning
model uh that I was mentioning earlier this is automatic we we don't configure anything and it's looking across every single column uh dynamically to see if there was a spike in NES empty string
zeros or drops in segments so this is a dynamic threshold sit on every column so one column it might be significant to have 5% NS that's abnormal so would report on that one another column that's
not a big deal but 75% NS would be a big deal um so drop in segment records is a special one where if I go to my run history I'm going to filter down to a fail check in the past so anoma will log
all previous runs and here what anom is showing me I'll just come down to this bottom visualization is in one particular column the merchant category code column we were just looking at it's detailing
out all the different segments uh so for example there's that online retail and food and we can see that the prediction intervals which are represented by Green that's what anom is expecting based on
the machine learning it's done uh now there are two segments in this data set for this one column that had significant drops in the number of Records uh so we could see fast food and fuel and if I
double click into fast food I can see that after doing the sampling the unsupervised modeling anomalo want to make extra sure that this really is a significant anomaly and it generated a Time Sur model on this
segment and it's showing that there was a significant drop yesterday we predicted to be somewhere around a thousand records but there are only 16 records that came in so I can't emphasize enough and this is a
completely different way of doing data quality monitoring I can't emphasize enough that these are completely automatic across every single column and the segments uh the these are looking
for Value distribution shifts and to make sure that previously unique columns don't have duplicate records uh the last thing I'll show is key metrics and validation rules so everything we've
gone through so far is completely automatic I mean that's amazing between all these visualizations and all these checks we're getting really good coverage out of the box but there's
still going to be a time and place for targeted checks so let's say I car about the average transaction amount you know I could be an analyst and you know I was talking earlier about self-service so
self-service I need this to be as easy as possible and anomo has a library of over a 100 out of the box checks um so really extensive most of them are completely no code so if I look for my
average I would just select my column mean check and from here I care about my transaction amount so I would just save it and run it and what that looks like I created this right before anomalous tracking that average
it generates a time service model and it detected an anomaly so I would have received an alert I can see there was a massive spike in the average that's a significant issue um so this is going to
be in a few clicks users can self-service uh so an analyst if they care about a certain metric where they have to report to their VP every week we need to make sure that's right so that
analyst would want to create their own monitoring metric so finally Val validation rules of that last piece for the critical data elements we want to make sure every single value is
following business logic uh so I set up a null check earlier uh and again this is no code same no code library to make sure every single value in a column is null not just a significant Spike and
ano found that 5% of the data in this column was null and if I scroll down a nomal will give me some sample bad records some sample good records and can even expose the SQL to copy and paste
this into Data bricks and retrieve those bad records um so I I I could grab all this and what's more anomal has an API and a python sck to use anything that's in anomo through that API so for example if we have a a
datab bricks workflow and we have a bronze layer uh a silver layer and right before we publish it up to our gold layer we want to run anomalo checks on the silver layer so we can say Okay
anomalo run your checks if things work we pass it on to the gold layer if not then we could pause that process grab the SQL statement automatically and quarantine those records out and not
only that when I click into that link when a a uh when an issue is found anomo will show me this automated recot cause analysis so this is completely automatic I didn't ask anomo to do anything and
what it found here was for these null records there was a significant Spike uh in I mean the those null records are coming from ride share transactions which are split between left and Uber so
in summary anomalo is going to be one first automatically giving it coverage from a data quality and table observability perspective uh with unsupervised modeling with um you know data volume data freshness and then
gives users that self-service capability uh and then finally it's going to be alerting and giving that root cause analysis um so that's all I wanted to show uh but I I want to emphasize that
this is going to be um displaying within Unity catalog as well for that tight integration I mean I'm amazed I mean it's simple and it doesn't require coding so believe everyone should be able to use
it and this and to be honest this last piece is show me like how you're able to write the SQL code that I can copy paste and bring again to to data braks to run I don't know query to get this code get
this data put it send it back in the quarantine in order to fix the issue and then send it back again to be analyzed and then push it it's really simple way to monitor data and make sure that
nothing messes up with the dashboard that's going to be brought to the Le sea levels or the uh head of I don't know company yeah I I'm I'm super glad this is resonating youf uh and I I I can't
emphasize enough that this you know now with all these llms and AI for doing analytics uh all of the sudden we're broadening the data that we could use there's a lot more data that we could
use with AI and because of that it's more important than ever to make sure that that data quality is going to be good everywhere and if we're not familiar with that new data we're going
to be using then we need a degree of automation that isn't just looking to make sure that new data is arriving and it's complete uh we need automation to detect if there's significant shifts um
and so that unsupervised modeling uh is a way to do that um so yeah really appreciate your time youf with that and you got my attention but next question what if I want to try analu is there in
trial period or how does oh yeah how could I forget the most important part yeah so we we offer free poc's um so feel free to reach out to us uh one thing to note is we are directly within partner
connect uh so if you head over to partner connect there's going to be the this big anomalous symbol right here uh click it and then you can get a free trial um so it's just a couple tables
what what I'd recommend is reaching out to us as well uh because we could do a more involved one uh where we can uh enable a lot more tables uh but if you want to give it a few clicks H go right
into partner connect amazing so I'll make sure to out the links in the description of the video and in case when I reach out I will add all the animal documentation below and I will also add the contact of
Zachary awesome all right well I I I really appreciate it Yousef thanks for all your time thank you thank you for being my guest and bye bye