The Struggles We Skipped: Data Engineering for the TikTok Generation

,

A tale of two junior data engineers.

Our generation of developers might have it “easy” due to there being a plethora of tools available to automate and plug and play everything. However, this abundance poses challenges in breaking into a field. This talk explores the perspectives of two junior data engineers—one entirely new to data and the other with a data science background—both navigating the complexities of data engineering.

The first one, a data scientist navigating her tasks without the luxury of well-formatted data. This journey inadvertently led to a gradual familiarity with complex tools like Spark, and the necessity of understanding various connectors and writing detailed code for data extraction and normalization. With the introduction of dlt, a significant shift occurred. This technology automated many of the tedious processes, allowing analysts to focus more on analytics, and less on tedious data handling.

The second one, never having had to deal with the chaos of unstructured data, was directly introduced to dlt. Spared by the typical struggles faced by traditional data engineers, she's set to find out what happens behind dlt’s automation throughout the talk. After realizing that the two lines of Python code she wrote saved her from the manual tasks of data normalization, structuring, and loading, she will gain an appreciation for the tools at her disposal, especially dlt.

dlt, or data load tool is an open-source python library for data teams of all sizes. It can extract a range of data formats from various sources, then normalizes that unstructured data into a relational structure and loads it into the destination of your choice. All of this is done within a few lines of Python code, as compared to the usage of different tools that were needed to get these tasks done. It is a valuable and cost effective addition to a company’s data stack.

The talk will follow a step-by-step, linear narrative to outline the challenges of building a data pipeline and illustrate how dlt can resolve these issues, thereby automating the process. Beginning with schema inference and evolution, then progressing to dependency handling and data governance, each challenge will be portrayed as a quest on the journey to constructing a well-defined data pipeline. As junior data engineers, we would like to emphasize the paradigm shift in data engineering towards a greater level of abstraction. This shift, enabled by tools such as dlt's declarative incremental loading, empowers junior engineers to tackle tasks that traditionally would not be considered junior-level work.

This session took place in track Data Handling & Engineering and was classified suitable for intermediate domain / intermediate python by the speaker.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:06]

Thank you. So yes, Gen Z has a different perspective on most things, so why not data engineering too? So anyway, thank you for everyone for coming. And a little bit about us. So we are working students at DLT Hub, and we come from, both of us come from different parts in Asia. I'm from Pakistan, Anun is from Mongolia, as I'm sure there are many different nationalities here too. So, yeah, we come with different perspectives on the topic that we're presenting, and you'll get to learn about us as we go through it. So, data engineering for the TikTok generation and the struggles we skipped. So, when I say the words TikTok generation, I'm sure there are many meanings in your mind as to what that might mean, and that's precisely the problem, right? There's too many of us. We come with, again, very different perspectives. We might be doing a little too much. And there's too much to keep track of. And to put it all into a single definition, it could be hard. But let's try. So, particularly, let's try and define the TikTok generation of coders. So, to make sense of it, let's think of what TikTok is for us. There's a lot of content. There's a lot to keep track of. It's flooded with information. It's flooded with new trends almost every other day. And it's really a metaphor for what we as a generation are facing, constant changes. And for those of us who don't identify with Gen Z or don't want to identify with Gen Z, you're still the audience. You're still facing this so to put it into our context we you know we are rap we are navigating a rapidly evolving tech landscape and we keep hearing words like learn fast build fast ship fast we even hear fail fast right and those things are incredible sure um and and keeping with those things we're also supposed to be building meaningful things but sometimes especially as we are learning We are learning how to create these new things, how to build these new things. It can be hard, right? Stuff like this happens. So you learn a new tool or framework, and then it doesn't stay new for very long. Eventually, another framework comes along, makes it obsolete, or, yeah, you know, it doesn't integrate with the other things that you build on top of with it. people who have worked with the first version of TensorFlow might be able to relate to the headache that it caused when the second version came around, and then some of the functions, you know, they changed the data types of what they were returning. It even was not compatible with the tensor objects themselves in the second version of TensorFlow. I love TensorFlow, but, you know, it hasn't been easy to keep up.

Speaker 2 [03:25]

So, yeah.

Speaker 1 [03:28]

Just like that. Moving on to where we are right now. So things aren't that bleak, right? Yes, we have a lot to keep up with. We have a lot to keep track of. But again, to stop being so bleak and to talk about one of the reasons that we can keep up, it's because of the preexisting tools libraries and modules, particularly in Python, that help us navigate through this craziness. It's stuff that you can build on top of. And I'm sure that for a lot of these different modules and libraries, we can't really imagine our lives without them today. And building these are the struggles that we skipped.

Speaker 2 [04:18]

However, this meme again, navigating this world full of tools available for every conceivable test there is, is the struggle we did not skip and unfortunately have to face. So to make your life a little bit easier, we're going to be talking about a tool that is going to make your life easier by challenging your data engineering challenges. That is DLT. So the agenda of this talk is that first Heba will share her experience with DLT as a new gen theatre professional, and I will then provide a general overview why it matters for us all, and then explain how it actually works. Now back to Heba.

Speaker 1 [05:10]

right so um i've done a couple of data roles in the past i really relate to this guy i've been a data analyst i've been a data scientist and you know the reason why i relate to this poor person is data analysts and scientists were usually in like this weird position of am i in the tech team No, am I in the product team, or am I in the business team? I'm somewhere in the middle, so when I'm in the party, I don't know who my friends are, you know? I don't know where my loyalties lie. So, yeah, that's me. But anyway, walking back into the different roles that I have served. So in the first one, I was a computer science researcher, which, if any of you have been a researcher or are in academia, you will pretty much understand that this means that I didn't my role really did not have a set definition I was doing computer science I was doing some data stuff with data modeling some analytic stuff some machine unsupervised machine learning I was doing some front-end development with react I was doing some back-end stuff with dangle I was doing some agent-based mode I was doing computer science. So, you know, whatever that means. But after that, it got a little better. I joined the industry, and then my role was, my first role in the industry was a business analyst, and that was at, I was a business analyst in operations at an expanding unicorn startup. And this was a fun one. The analytics was fun, the dashboarding was fun to the point that, you know, you would forget that their dashboards were on Google Sheets. But it was nice. After that, it was a data analyst. It's a bit of a different role. This was another startup. It was, you know, very young, fun, high energy, all those startupy things. It was really fun. But a lot of the times, my dashboards, the ones that I built, used to be sitting on top of broken data pipelines. And to illustrate what that feels like, imagine being in an all-hands meeting, right? All-hands meeting. And the dashboards that you built that have been working fine for perhaps the last couple of weeks, they've suddenly gone bonkers. And the reason for that is not in something that you've done, but because it's sitting on top of a broken data pipeline. And now I am a data science working student at DLT Hub. And I'm finally learning how to get into the world of data engineering because clearly I was facing issues in that regard. And I'm learning how to fix the previous problem of broken data pipelines and the feeling that comes along with it. So to understand these problems a little bit more, let's get a little bit more technical. And let's start with ETL. So extract, transform, load, extracting data from sources, transforming them into some sort of a form that you require it to be in, and then dumping it into some sort of a destination that you require. It's a process that not just data engineers have to take care of. It's something that, as a data analyst or data scientist, you need to know how to do, be aware of it, because it's literally part of your job to be able to take out insights from anywhere. And not everyone is kind enough to give you a working ETL pipeline. So with that, there are some issues that people like me have to face that I'm sure you guys will be able to relate to. Yes, of course, starting with JSON strings in the database. Then there's ad hoc sources, right? So the data can start flowing in from different endpoints and it can look like different things. perhaps your tech team is working on something new, they wanna test it out, and you have to plug that in, which kind of leads to ad hoc analysis, you have your own analysis projects going on, but then stuff like data requests come in, about 20 a week. After that, because again, we're not very lucky, we sometimes get unstructured data because it's not a guarantee that you'll get it in some sort of a table or a queryable object that you would need it to be in. And with that, we have ad hoc analysis of unstructured data, the entire Japan. And because of these, you know, what seem like little problems, but they're not, analytics tasks become quite heavy ETL tasks before anything else because of these things. And then on top of that, you add problems like finding the right tools for streaming that take you away from your home as an analyst, which is either Python or SQL. You don't wanna do that. Then again, on the topic of learning tools that you don't want to, to fix problems that you don't want to but have to, there's stuff like learning Spark on the fly. So yeah, there's a lot of stuff that you might not want to do but you would have to and it would be nice to be able to encapsulate that somewhere. So of course, as I said, it takes time for us to navigate through these problems. So with that, let's come to time to data. So it takes time when you have specially unstructured data. For example, you've got something like this. This is one of the data sets that I work with. And it was nested, it was madly nested. And the thing with something like this is you have functions and operations in Python that can take care of unnesting and structuring your data. But when you have nesting on different levels, so for example, you see here that this nesting is not just adjacent within adjacent, it's adjacent and then there's a list in it and then there are other JSON items inside of it. So this translates to a whole other data model, right? That's not something that Pandas promises you to do. So in that case, you would first have to spend time on picturing, of course, after exploring the data, you have to picture how you would want that to look in your end product and then maybe work your way to some sort of a model. And do it fast because no one cares about your ETL. They want the analysis, they want the insights, right? So, it can, the solution can only partially look like this, right? This was just some of the stuff that I did to try and get it into one table. However, this is what it's supposed to look like, right? It's what a certain magical, you know, library in Python can do. So, yeah, it's, and everyone who works with data, you know, if it works for you in that situation, a relational model is something that we really look forward to, right? It helps either in your IPython notebooks or it helps in your dashboarding. And then there's stuff like ad hoc data sources that come up again because I love them, right? Whenever you have to fix for those problems, you have to look into different connectors. And those connectors come with their own different learning curves and they obviously come with different prices as well that you have to justify to your team. So yeah, it's a whole bunch of problems that we don't want to face. And then enter DLT. So it would be pretty cool if you can skip over the stuff that we just went through and then take your raw data, plug it into some sort of a magical API and then jump directly into analysis, which is what your job is supposed to be. And so when I envision some sort of a plug and play solution, it hopefully does not look like this. So and with that, we come to a little bit more of exploring what DLT is. So for analysts or actually all Python users, whenever we think of the plug and plug and play, it usually starts with a pip install. So going back to how helpful these libraries and modules have made our lives, thanks to the people who came before us who built them, we don't really have to worry about writing nested loops to filter a table. We just use pandas. We don't really have to write loops to normalize a a list, we just use NumPy. And just like that, DLT or data load tool is my pip install solution. It's an open source Python library to build and automate data pipelines and structure your unstructured data. Which is my favorite part, if you can't tell. So with that, let's jump into this plug and play paradise. And this particular example that we'll be looking at is a very special part of the DLT Hub office. It's this banana light right here. It's right over our CEO's chair among many other banana-themed things. Don't know why. But, yeah, one day when he wasn't in the office, we connected it to a smart device, a smart plug, and then we connected the smart plug to our Python script, and we wanted to store the data coming from this light to a destination which for us was DuckDB. So let's see how this process looks like with DLT. So the first step, we declare a pipeline, right? The pipeline, we declare that with a name and a destination. So for us, as I said, the destination was DuckDB, but changing the destination is just as easy as just changing the string as you can see here. So that's the first step, right, declare a pipeline. After that, we come to declaring resources. Now what the actual definition of resources is is something that Anun is going to go over. But for now, we can just think of this as we are getting data from different endpoints, you know, ad hoc sources, different endpoints, and these are all data packets of different sizes. And what we've done here is we've taken status, specifications, and properties, and we've declared that this is what's coming into the pipeline, right? And after that, the last step is basically pipeline.run, and then we're running all the resources that we declared in the last slide. So we declared the status, the specifications, and the properties. You know, they're all yielding different endpoints and different data packets, and the pipeline is running them, right? And just to highlight, our pipeline is up and running, by the way, in that last step. Just to highlight how, when I said that, you know, there's raw data and then you jump directly into analysis, this is what you see on the screen, is basically the development environment that DLT has with Streamlit. So what I've done is I've directly queried the active states for the bulb in the last some sort of a timestamp. And it's just as easy as plugging in your raw data and then getting into some sort of an analysis form. And there we have it. That was a very low effort plug and play solution with DLT. And that's it for me.

Speaker 2 [17:31]

Okay, you guys gotta wait, because what if you're a total newbie? And we've probably all been there, and some of us are, but okay, I need to go beyond my prepared speech, because I'm seeing you all guys and realizing that there might not be a lot of juniors here, but it might be just the fact that juniors also look like veterans in this field, because, you know, work stress, but, yeah, but I'm admitting that I'm a newbie, and I'm relatively young, I'm not sure if it's a good thing or a bad thing in this day and age, have very little experience, have some, you know, confidence issues, and make dumb mistakes, like the fourth point on the slide, yeah.

Speaker 1 [18:21]

Thank you.

Speaker 2 [18:22]

So, I want to start from the beginning and talk about what data engineering is at all. And as Heba mentioned, it's about maintaining ETL slash ELT processes, right? That is pipelines. And as a random data engineer on Reddit wrote, data engineering is like being a plumber. engineers pump data from one place to another and make sure that it's clean for the business to use. So in other words, it's the job not many people want to be involved in. And not to offend data scientists here with this meme, because there's a plot twist, right? Data scientists deal with data engineering jobs, too. And my co-speaker can vouch for this. And the reality is that data engineering is just unavoidable. Every data job is a data engineering job. But the problem is that not everybody is trained to be one, or simply not paid enough to be one along their main role, right? So this is why we're talking about DLT, a pip install solution that is open source, has many integrations, and is naturally Pythonic. And it's basically an in-house data engineer in a box from a very not so vantage but still very valid junior point of view. I think DLT is great because since it's open source, there's an open community where we can learn from seniors. And obviously you can take up some tasks to enhance your skills. And since it has many integrations, regularly tested integrations of sources and destinations, It saves you time for other things, say data science. And naturally it's Pythonic, so it's something that comes naturally to us. So there's no steep learning curve. And yeah, there might be seniors sitting here, and I'm seeing many of you and thinking, okay, maybe you're just probably enjoying the memes, right? But yeah, DLT can also make your life easier, because since it's open source, it's cost effective because saving money matters when you're juggling more important things. And since I mentioned it has many integrations, juniors save time, which means there will be less excuses from them. But jokes aside, there's just no need to reinvent the wheel. Maybe it's just that you don't have junior data engineers. You're the only one, maybe. And since it's Pythonic, it runs where Python runs, which is basically everywhere in the the data world. So, now we're done with the general stuff. I want to explain how it actually works and is making your life easier. And we got to go back to Heba's code demo. So, in the first step, we declare a pipeline. And what's basically happening is we are declaratively creating a connection that moves data to a specified destination. This is your data pipeline. That's it. And what you've got to do now is just to define the data objects you need to pass to this pipeline. And in this case, in step two, we define three DLT resources. And you might be thinking, what are DLT resources, right? And then they are a logical grouping of data of similar structure and origin. So they basically represent a table in a data set. They're defined as decorator functions that generate data on the fly because storing data in memory isn't a good idea. And they're eventually passed to the pipeline to be loaded, for the data to be loaded at the destination. And what's great about DLT is that it supports incremental loading. So in the data object, the default write disposition is append, but if you need to do deduplication or upsetting data, you just gotta set it to merge and also provide a primary key, and then everything is done. There's also a replace option, which is quite self-explanatory, I think. So, in step three, we finally run the data. I mean pipeline, sorry. And alternatively, we can run the pipeline only once by defining a single DLT source. And our DLT source is our logical grouping of multiple DLT resources. Basically, source is a data set, and resource is a table. That would be the easiest explanation. But before we pass that source to the pipeline, we've got to define it right. So we go back to step two. And this is what we had. And this is how we define the source, and eventually run the pipeline. OK, now you might be thinking, what's the point of all this grouping resources into a source, right? And imagine you have 100 different API endpoints, each representing a separate table. And by grouping them into a single source brings efficiency because you run the pipeline once, meaning you load everything all at once, while having control because you can access resources from the source and apply transformations from the source again. And another advantage is reusability. So you can dynamically create resources. Say you have many endpoints, right? And they probably require similar authentication, pagination, and data extraction methods. So you define a single method, like the getResource method here, and you dynamically create everything else. So in the end, you can view your data in QueryAd, And there's a pleasant surprise. You see that the nested data was separated into parent tables and child tables. And what happened behind the scenes is that DLT recursively unpacked and normalized data into relational tables. And you didn't have to do anything except for declaring the pipeline. So this is the install solution. Yeah. Just something that happens automatically, so you can chill, right? And honestly, I must admit that I was introduced to DLT straight away without having dealt with data that much. I didn't really appreciate it that much. Then I tried to replicate what DLT does without DLT, and then I spectacularly failed. Yeah, that happened. So try out our demo in our GitHub repo. There are also a lot of other demos with all kinds of integrations, don't hesitate to contact us. And obviously, visit our website. There's also a GPT-4 assistant. So yeah. Thank you.

Speaker 3 [25:34]

Yeah, thank you so much, Hiba and Anun, for this great presentation. We have a couple of questions. I'll start with the question, how does DLT scale? I assume the meaning with the size of data and the quantity of data.

Speaker 1 [25:52]

We have a bunch of companies using it right now. However, I would say that if you want to learn more, you can go to our Slack to see the use cases that are applied.

Speaker 2 [26:03]

Yeah, DLT also supports has async functions so you can you know asynchronously run multiple

Speaker 3 [26:13]

I think, Hiba, you mentioned at some point a magic library that lets you extract the table structure from a nested data frame, and the audience was curious about that.

Speaker 1 [26:24]

It was DLT.

Speaker 3 [26:27]

Yes, and would you like to maybe compare it to industry standards, the AGD tools like Airflow?

Speaker 2 [26:42]

Okay, we've got to say that we're not competitors of Airflow because Airflow is an orchestration tool and we are a data loading tool. And I want to say that DLT works with Airflow and Daxter, that is. I tried it myself.

Speaker 1 [26:59]

Um, can

Speaker 3 [27:01]

Can one use SQL for the transformation part in the ELT process? And does it work well with relational database as source and destination of the pipelines?

Speaker 1 [27:15]

So, again, right, DLT handles the loading part, and it's pretty cool if you can patch it up with DBT for transformation, and that's where you could perhaps use DBT both in Python and on as SQL models as well for the transformation. As for the relational database part, I didn't understand if the data already exists, of course it will get passed through and loaded into the destination as a relational database. And if it's not, if it's something like some very nested unstructured data, then DLT handles that structuring and it creates parent-child relationships between tables and then it does become relational.

Speaker 3 [28:07]

I think, last question. I think in the demo, the data you presented, the data was automatically unnested and normalized at the end. How much control does the user have over these processes?

Speaker 1 [28:21]

Yeah, so you can, at the part of declaring the pipeline, you can also give it the schema as a YAML file. So if you want DLT to follow that, you have some sort of control over the schema. And as for the unnesting and the normalizing, I suppose that is for what DLT handles. But for the schema changes, you can definitely control that.

Speaker 2 [28:55]

You can also specify the level of unnesting.

Speaker 3 [29:00]

Okay, let's thank again Hiba and Anon.

Anuun

About — in the speaker's own words

Writer by choice and a data enthusiast at heart. Crafting compelling narratives with Open Source Software at dltHub. With a background in International Relations, I am currently pursuing Computer Science, focusing on Machine Learning, at TU Berlin.

Hiba Jamal

About — in the speaker's own words

The data field has been my home for 3 years. I'm now a Data Science Working Student at dltHub in Berlin. Previously, I contributed as a researcher, data scientist and business analyst in startups and government-funded projects in Pakistan. Currently pursuing a master's degree in data analytics and AI for business management, I hold a prior degree in Computer Science with a touch of liberal arts.

Social card for talk: The Struggles We Skipped: Data Engineering for the TikTok Generation