How to teach NLP to a newbie & get them started on their first project

The materials presented at this tutorial were initially created for high school and university students to help them to get started with their first machine learning project using textual data. Machine learning on textual data is more accessible for beginners because it does not involve missing data imputation, normalisation and scaling. It is also easier to analyse and interpret the results (e.g. why something was misclassified). There are many introductory courses on NLP on the internet, however, they are not for free and they either only cover complete basics¹, or do not cover machine learning algorithms² and treat models as a black box. Also, they do not show how to do research correctly (e.g. setting a baseline, making design decisions based on correct validation etc). These materials in the form of jupyter notebooks can be used by teachers to guide their students through an NLP research project from start to finish.

These materials are of course not limited to teachers and tutors at academic institutions. Many companies rely on customer reviews, social media, client records, and various other content created in natural language, but often use sub-optimal solutions to analyse it (like MS Excel). These materials will give working professionals all the tools to get started with text analysis, as well as teach them the fundamentals of machine learning, so they can automate document labelling and other manual tasks with the help of document classification (e.g. Is a customer review positive or negative? Is a certain document about topic X or topic Y?). A minimal understanding of programming (in any language) is required. However, all necessary Python libraries will be covered.

The aim of the tutorial would be to present the materials which contains 7 “lectures”, several practical exercises with solutions, and a case study and hence can be covered in either 10 hours (10 weeks) over a term or a 2-day workshop.

¹https://www.udemy.com/course/natural-language-processing/

²https://www.udemy.com/course/nlp-natural-language-processing-with-python/

This session took place in track Natural Language Processing and was classified suitable for intermediate domain / intermediate python by the speaker.

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:04]

First I'll talk a bit about myself and who I am and why you should actually like trust me with the materials that I'm presenting And then I will basically present it's not a live coding session Basically, it's already like materials that I'm using to teach university students and working professionals Who you know want to like kind of like, you know break into like machine learning and you know traditional NLP and and then yeah, as I said, I will go through the through the notebooks and just talk about my experience and and, yeah, and, like, how to best, you know, use those materials for teaching. So my name is Lisa Schadler-Green. I'm originally from Belarus, but I grew up in Germany, but I've been living in the UK for the past 11 years. I did a PhD at UCL in computer science, to be specific, in computational argumentation. And during that time, I did a lot of teaching. That's why my PhD took a bit longer, not because I'm lazy, but because I did a lot of teaching, and I was very good at it. so um however academia doesn't pay really well so i left academia but i still enjoy teaching and like doing it you know as kind of like as a side hustle so um i have private clients it's as i said either working professionals um or university students who need like help with their courseworks or mentorship for example um um i did my bachelor was in law and then i switched to computer science right so i did one of those conversion courses and so quite often i have people like students reaching out to me being like oh yeah i'm in a similar situation like i want like you know a mentor that i can talk to about courseworks you know get your opinion on stuff and or yeah just help in in certain subjects and so this is also how those like materials came about right i created them for students and and then eventually i got a position as a as a guest lecturer at the University of Oxford in the Department for Continuing Education and so there I run like small boot camps like weekend courses and this is basically like what it's tailored for right it's like it's a good it's materials for to cover in two days it had really good reviews I was launched it last year and this year for example it's completely booked out so yeah that's that's me in a nutshell and why I'm doing this another motivation for this was actually my brother who worked as a data analyst in the public sector, and they were using Excel to extract data from text. Classic scenario, right? Which is ridiculous. And I know there's a huge hype now for generative AI and chat GPT, et cetera, but there are also a lot of people out there who don't need that, who basically just need to learn how to you know use regex or traditional ml methods or what is also really popular which i will not be covering here but is for example named entity extraction to know how if you already have a dictionary of words how how you can train your own name entity classifier so things like that so that's what it's tailored for no gp chat gpt today no deep neural networks just all good traditional machine learning using using scikit-learn and so yeah so for those who just joined i will be there will be no live coding it's uh i'm presenting material teaching materials that are kind of like tailored as a two-day course as a two-day boot camp for people who are just getting started with machine learning and nlp um so yeah as you see it has like a folder structure like day one and then there's a day two and so the very first lecture as i said like i usually talk a bit about myself and why you should trust me and then i talk about you know what is nlp right so um i stole this one actually from my from the lecture slides of my favorite lecture at ucl in nlp sebastian riedel maybe you've heard of him so i really like this one where he's like okay no nlp is not control f anymore i'm also not claiming that but this is where it started right and and still you you know like still a lot of problems can be solved you know in a way easier way instead of just you know throwing a transformer model on them and this is kind of like the philosophy i'm representing that you know first try simple way before you before you do it like the very complicated hard way then we talk about applications in lp usually i make this like a more interactive session where i ask people you know where do you think you've used nlp before um you know obviously everybody has used machine translation at some point in their lives or well they know how spam detection works so yeah so we kind of have like a group discussion and then we talk about challenges why nlp is so challenging well because the data is unstructured right it's not like you have just you know beautiful columns of numerical data um then obviously there's ambiguity i mean some of these things they apply more for generative AI but still and then we so we talk about like some examples so again credit to Sebastian Riedel for this one you can't eat a dumpling wearing a tuxedo well but the dumpling is not wearing a tuxedo so you know like I'm just trying especially you know in the UK the UK is really international I have like people who who's like where English is their third or fourth language so I'm trying like to make it a bit like easier digestible and then we have lexical ambiguity that we talk about structural ambiguity and then I like some really nice examples so I really like this one it was down for a while but it seems like they put it back up so what we usually do is we play a little game it's called ambiguity game not sure whether you've you've heard of it before and so you read the instructions so basically yeah but as I said so I will I will let you click through it yourself but so here that's basically the example that they use I can't remember from which university that is there's also a paper here that you can read and it's called madly ambiguous and it's from the University of Washington IBM okay and and so basically the idea is to to make people understand you know that although structurally a sentence is the same I'd saying for example Jane ate spaghetti with a fork and Jane ate spaghetti with meatballs it's it's it's not like meatballs are not not a utensil right so this is like kind of fun and interactive and I really like this one, so I'm really happy that they put it back up. And then you can basically play, so here they explain it, right, what they like about, they talk about syntactic structures, and then you can play a game. Finally, at some point, come on, there we are, begin, right, and then you like, you ask students to to engage right and then um which one i really liked is jane ate spaghetti with lady fingers lady fingers are cookies not lady fingers as little fingers of a lady and then yeah see and it doesn't get it right because he thinks that you're like like it's the fingers of a lady but it's actually the biscuits that you use for tiramisu so so yeah so that's fun um Then we talk about synonyms, and given I presented in the UK, I use British English as an example, what Brits say and then what they actually mean. I will not read them out loud, but you get it. Then there are some really beautiful demos from Allen NLP that you can check out, I'm I'm not going to go over them yet, but I included some screenshots where you can, you know, write a sentence yourself or like a passage of text and then question it and see whether they got it, whether they get it right or not. Then, again, what is another challenge of NLP is actually just a change of language, right? So especially in generative models, that's a big thing. So a couple of years ago, you know, nobody would know what a PCR test is, right? Like because there was just this data was not there. Like all those big language models, right? they're trained on on corpora of data and and yeah so I like this one so in the past we drank way more wine and had more fun and now everybody's like super health conscious right and it's drinking water so that is the Google n-grams view it's really fun so obviously you can try to come up with some other examples so I like this one maybe don't use it when you teach it to under 18 and then And figurative speech, another demo from LNNLP, where you use sarcasm and see whether the model gets it right or not. They must have trained it, because when I just started using the example a bit over a year ago, they would think it's a positive sentence, and now he's somewhat confident that it's a negative one. I think I just used the example too often, and now they trade the model in the background. So this is the first icebreaker lecture that I give to students, as I said, either in a group or privately. In a group, obviously, it's more fun. Cool. Then what I do is, so I do tell people, right, that as a prerequisite, so if it's university students, not a problem anyway, because usually they already know Python. If it's people who sign up for the course that I teach at university, then I tell them that they should do at least some basic course, just to know the syntax. I don't need them to know perfect object-oriented programming and how to write classes and whatever, but at least they should be comfortable with the syntax. Otherwise, because it's very code-heavy, I don't want to give theoretical lectures where people kind of go out and say, okay, I got it, but I have no clue how to implement it. So I give them a lot of code that works and where they can just use their own data and it will compile. So yeah, so if you don't know, if you've never seen Python code before, probably it's a bit intimidating and challenging um but basically i do still spend like 40 minutes i will not do that with you guys now though um to just go through some very basic stuff and um yeah and obviously it's up to you whether you even include that right you can just say hey no i don't have that much time the prerequisite is you know python more or less before you before you do this um do this course cool so that's that then i also have then usually you once we're done with this to again make sure that people weren't just nodding and saying yeah yeah sure sounds good yeah yeah um i um i tell them now you have an hour to do a practical right so i've created a practical it also has a solution and then depends whether somebody usually i like walking around because it's usually it's not more than 20 people and then there's no point like then doing it together usually just walk around and then check in with people um and see whether like how they're getting along and it's very like simple you know but a bit it already involves a bit of text pre-processing right they have to read in a file they set it to lowercase the lethal punctuation and then this also gives you when you actually walk around and give this like in a classroom setting it also gives you an idea of like you know the level of of people yeah then there's a solution included that people then usually can just have a look at and then we start talking about text pre-processing and so this is where you know I'm now you know gonna give you like some examples that I've noticed that in other NLP courses online like certain things were missing or were omitted that I then then tried to you know like try to include in my well in this in this little course that I've created so for example first we start talking about regular expressions yes boring but sometimes does solve problems to give you an example because I have a degree in law I sometimes take projects on and read in legal tech and if you've ever seen you regulations like the titles are freaking massive and they so they have abbreviations right and then you could either go the fancy way and say oh let's train you know they can then named entity recognizer for that or you can actually just figure out and write a good regular expression for it and so this is why I like also talking you know about that i'm not telling people should now become like experts at writing regular expressions but they need to be aware that something like that exists so that they can google it right um also in a second i'll talk about text normalization where again like this becomes relevant so i just tell them yeah look just here are the tables this is what you can do with it then run through a few examples of you know how to apply uh regex in in python and that's it then i tell them you know most likely you know at the beginning you will probably you do like Kaggle challenges or, you know, your textual data will come actually in an ICSV file, so you should be able to at least, you know, know some basic Pandas. Again, I'm not giving, like, a huge Pandas tutorial here. I just say, okay, you can go away and learn it by yourself, but at least you should be able to, like, you know, read in a file, select some columns, you know, apply a function to a column, etc. So then I go over that. So nothing super special. And then we start talking about text pre-processing and so this is what I've noticed is like very often if not always omitted and I also see it in a lot of research papers that people do it wrong and we will talk about this in a bit when we start talking about baselines that people just randomly apply pre-processing without ever verifying whether it actually does any good to your data so it's like oh we set the baseline we already removed stop words okay have you actually verified that that did something good is it actually useful to remove stop words um so we'll talk about that in in a bit and so yeah so basically especially it's more targeted towards you know people who you know who are new to this and who think that there's like one way of doing it you know well there's not it obviously depends on your textual data whether certain certain decisions make sense or not and so I use like three different examples here here we have you know like a spam text a review sample and a new sample right and different types of text require different pre-processing it's for example spam or twitter data is obviously very dirty yeah um then you have like a new sample you know which is like perfect nice english with nothing nothing funny going on here and so um this is where you know i kind of like gap the um well the the the bridge to uh text normalization and how i call it and like i've seen it in some other um like some other books using it as well So, for example, if we now start deleting punctuation, for example, from a spam message, right, we get a lot of random stuff flying in here, right? We have, like, the HTTP, then we can see, okay, we deleted some punctuation, we have the www, then we have, like, you know, it could be, you know, the domain name could be unique, it could never come up in the data set again, so it would not really be a distinctive feature, right? So this is why I'm showing, you know, think first before you apply some text pre-processing. So what would be a better way of doing it? So here I give two other examples. So for example, in a review we might also lose a bit of information, we lose the emojis, whereas in a new sample it doesn't matter whether we delete punctuation or not. So what would be a smarter way of dealing with it is actually to normalize the text, just how you normalize numerical data, and you can obviously also normalize normalize text so you can find patterns right like email addresses HTTP addresses money symbols you know you name it right obviously depends on your data and then use regex to actually normalize it and this way you it's you actually already generate features right that's a part of feature generation that now instead of having different domain names in every you know in every spam text now you just have you know email address or web address, right? So very important. Then we talk about tokenization, obviously in the English language, it's rather straightforward. I introduced students to the NLTK package, which already has, you know, a word tokenizer in there because the, like, you know, the basic Python split is naive, right? It keeps the full stop at the end of a word, whereas NLTK is a word tokenize. Another nice thing that I think I've never seen in a paper that actually uses Twitter data is NLTK tokenize actually has a tweet tokenizer, which actually keeps emojis and hashtags and stuff together. So that's another thing, right? And this just comes with experience. I think I've never seen this covered anywhere before. I actually myself stumbled upon it a couple of years ago by chance. So also really nice, right? Because this way, actually, again, you don't lose features and you don't lose information. Then we talk about stopword removal, which again, you know, you need to experiment with it. You can't just say, oh, I pre-processed my text, I'll remove stopwords. You need to verify, you know, using a, you know, a validation set whether any of your pre-processing actually makes sense. You can't just, you know, suck it out of your fingers and say, okay, that seems logical to me, I'll just use it. You need to verify it. Again, as I said, will be covered in the next notebook. And then, yeah, obviously also which stopwords list to use. It's often you need generate your own special stop word list so again when i was doing a project in legal tech and we were dealing with regularizations obviously the stop word list we used was you know was customized right because all those regulations had like very specific uh legal lingo um like for example the word paragraph that we just wanted to to get rid of so sometimes you know again you can't just use something that was already created and is out there sometimes you need to do it yourself then uh lemmatizing and stemming so um i talk about the differences between you know using nltk and using spacey because nltk is actually more annoying for lemmatizing because you also need to provide the information what type of word it is so it only by default works with nouns so again you know you can sometimes tell when you read the research paper and go into the code that they didn't know that and they actually didn't properly lemmatize if they use lemmatization So either use stemming, again, verify what works better, whether there's a difference. If you want to lemmatize, use spaCy because there you don't have that problem. Cool. And then after talking all those pre-processing, then, you know, we start talking about how to then, once you've pre-processed your text, how to actually then create numerical data out of it that you can feed into machine learning model. So for those who are not familiar with the bag of words model, basically uh how it works is that you have a bag of words which is all your vocabulary in your um in your corpus and then you create one hot encoded vectors right so if this is our vocabulary then the word they will be encoded at the first position we will be encoded at the second position etc yeah and then you can use either min pooling sorry some pooling or max pooling you know to either just you know basically tick whether a word is present or not or whether you also want to to track um the the number of times it's present again need to verify which one is better don't just use one and say okay that was my decision without actually basing it on some underlying evidence cool then a few other things which again come with experience or if you actually very good at reading documentation and not just skimming over it so also something that I've realized way later than I would like to admit is that when you use then the count vectorizer the IDF vectorizer from scikit-learn that it actually deletes punctuation and deletes single characters so again if people actually thought they did some cool pre-processing thought they actually kept punctuation and then they blindly use you know just the count vectorizer without feeding in any parameters they actually deleted it all so again read documentation and so yeah so this is what i show here right to see look you need to include the the token pattern to make sure that it's not deleted otherwise psychic learn deletes it then we talk about the limitations of the bag of words model obviously the bag of words model does not track the well the order of words so two you know opposite sentences would could actually mean then in the bag of words model would be represented the same way and we cannot take uh new words into account right if it wasn't in our training corpus then we don't have that information like there's no way we can learn that word so it will just not be accounted for cool so that's intro to text pre-processing and the bag of words model and then again i break the lesson usually then i think it's kind of lunch time and after that i ask people to write their own pre-processing function which basically you know it's just teaching them how to use those notebooks because there's nothing super new here but I want them then to engage with the notebook and find the information that they need and they can just copy paste it and create their pre-processing function right and then it also has a solution which again it's up to you whether you walk through it or just let people then figure it out themselves and then the fun stuff starts so supervised machine learning 101 right so I focus in this very first you know breakthrough into NLP I use classification as as an example, because then you can also use, you know, well, it to do something like that, like it's also introduced them just to some fundamental machine learning and good practices. So I talk about what what ML is a few fun graphs about the different types of machine learning about well, about the expectations and what we actually do as data scientists. yeah then about bias and variance so you know all the fun stuff so for regression versus classification so in NLP we are dealing with classification problems and then how to deal with missing data which in NLP kind of is well very different into numerical data but I still like to talk about it but in NLP I tell students well if you kind of just have a tweet missing well obviously there's just no way you can infer it but if you actually want to for example generate more data what you could do is you could translate it to another language and translate it back that's like one way of for example generating more more data if you have like one group which is um underrepresented outliers again not really a thing in uh in like in like in the specific you know like kind of tweet or text classification problem but again you know i'm talking about it because this kind of is not really nlp related it's just machine learning 101 then how to validate stuff correctly right so that you need a training set a validation set and a test set if you don't have enough data then maybe like don't have a validation set just use cross validation what cross validation is little graph here right so again i'm not going to a lot of detail here because i'm sure you've all like heard about this before different evaluation metrics you know why when you choose accuracy when to use f1 score what kind of data words where is it more important to use precision where is it more important to focus on recall things like that and then my favorite one which literally is just never taught and maybe you need to do a phd for that i don't know but basically how to set a baseline correctly so this is what i was talking about earlier right when i was talking about like all those pre-processing stuff that like people just very rarely set a baseline correctly if they like base their work on another paper that's fine right they say okay in that paper they did this and they achieve that but if you actually have a new data set and especially like in research right if you you know you're doing a phd or you work in a research group you actually collected some data and you even like paid money and made people label it and now you want to work with this data set and you want to do something with it then you need to set it correctly right because you have nothing else to compare it with and um and so like to use the example from previously your baseline then should be using the raw text right it's like okay using the raw text and some specific you know machine learning model classifier whatever i get baseline x and then you start you know iteratively you know figuring out okay should i delete punctuation should i you know generate features should i do this and always verifying with a validation set right like setting the baseline aside and then comparing it at the very end whether it actually made sense to to do all that so um so yeah so i focus on that a lot to really like make people understand that and then again old good scikit learn you know different you know to the label encoder count vectorizer we already talked about checking how balanced the data is you know checking for missing data so you see this is what i meant i mean it has a lot of code but you know you can literally just load in another data set and you know it should all still work more or less you know obviously like changing the the column names and things like that and how to use pipelines in scikit-learn um because and then like what i talk about is i'm not actually sure whether it's it's in this notebook is about uh data leakage right especially in the bag of words model so you and in second and pipelines they nicely they they take care of that right so that you don't actually learn the whole vocabulary and then start splitting right when you actually already cheated because you you learned the whole vocabulary so i make sure that people you know know what a pipeline is know why they should use it and not just use everything separately and then apply it one by one cool yeah how to set the baseline for example with naive base maybe to start with right how to interpret a confusion matrix how to print out the classification report how to use different how to use different models that and psychic learner just super easy and nice nicely to import short introduction of what each one is how to use cross validation so because we set the baseline right we set it aside that piece of data and now we're using cross validation and you know what the what the multi-layer perceptron is so and then you know the exercise would be if there's still time you know just you know play around with it compare all those different models and and make a decision and compare it to the baseline did it perform better or not cool so that concludes day one um so right so this is like a sample sorry i forgot to to to show it at the at the beginning so this would be like you know like a two-day program right so we would have what is nlp introduction to python um the practical then lunch getting started another practical machine learning fundamentals and that kind of like concludes the day for that day and then because it's the uk people just go for drinks and forget everything they learned um then day two so day two um at least like the way i do it is that um because you know if people got too drunk the night before then nobody's there at 10 a.m straight so you start with a like with an exercise again and just see whether people, you know, eventually show up or not, because nobody is there at nine or ten straight. So you just say, okay, now based on what you've learned the day before, write your own classifier, right? So, you know, use a data set. It doesn't have to be this one. This is just an example. And then set up a baseline. We didn't, like, use the count vectorizer, experiment within, like, the n-gram range and, like, other parameters of the count vectorizer. for example um there is um there's a like a minimum count that you can set for words this way you can for example get rid of all unique tokens because unique tokens don't um don't contain any any value right because they are not distinctive features so you can already start cleaning uh with the help of the of the count vectorizer and you don't have to do it yourself using you know collections and counter and stuff like that yeah um there is no solution for this because it's all going to be covered in a case study which usually then is the live coding session i just ask people you know did anybody of you find a cool data set i give them like a bit of time um or they do it already here right while they do the practical and this is done where um where there is like a live coding session then um what i do is i actually talk in depth about two machine learning algorithms so that it's not just which again is a problem what i've encountered with many of those online courses is that it kind of treats it as a black box it's like yeah just import knn just import mlp just import whatever but they actually don't really go into depth of any of those algorithms or it's the other extreme and it's just too much maths and you know and then and then people again get scared off so i want like to strike a good balance so now already you know at least how to import stuff and um and now let's actually talk about um two algorithms that are more or less straightforward to understand especially knife base in more detail so here i took as i said i took a knife base as an example because it's just very very easy to explain using an example and just basic maths from high school right so i talked to them about okay let's assume you have a corpus of data um obviously in this case it's a really really small one like um yeah what great match the election results will be out tomorrow the match was very boring it was a close election so label sports not sports sports not sports then we make it even easier and say hey let's assume you removed stop words and this this is your this is your corpus right it just contains a couple of words and now how do we train an if base so prior we have two sports to non-sports articles then you know how to get the likelihood right um how to calculate that and then we talk about you know okay but what if a word actually isn't present then you know your likelihood becomes zero so this is where smoothing comes in so i um i show them you know because because like i want people again obviously in psychic learn this already happens in the background like they will add smoothing for you but i want people to you know to understand that okay you should really understand an algorithm and not just blindly import stuff and because then you will never be able to debug it if you just don't understand how something works so for example naive base right is the prior is too small yeah it doesn't really work well for for imbalanced data sets right so again so just to to show people that um yeah please don't treat everything as a black box like a lot of courses do and then the complete example like the the maths and then the same implementation in psychic learn for naive base then i do the same for logistic regression i apologize for those really ugly pictures i did it on my old tablet i still need to update them but basically then i talk about you know the different um how like what a um what a word embed well we'll talk about word embeddings in a bit but um what a word vector is and then how word vectors are kind of like represented and how we need to learn the weights in order to then get the score in order to you know calculate whether something belongs to one group or the other talk about show them like a graphical representation of what we're interested in right to create the to create the hyperplane that splits one group from the other so again what the loss function is because again in naive base you can't really you well you don't really well you don't have a loss function and i do have another materials for an ml course where i actually go into depth about a gradient descent here i'm kind of skipping it but i still don't want to like completely ignore it so then so here i'm saying okay so you see for naive base it's just a calculation that's why naive base is actually also really fast there is actually or machine learning happening in there whereas logistic regression is a different scenario you have a loss function you need to minimize the loss and then you use an algorithm called gradient descent in order to you know continue minimizing it until you find the optimal minimum so this is like again a little you know teaser on on actual machine learning and then people can go away and do more reading and more you know like self-study on that an example how to use logistic regression against using using the same example yeah and then how to actually analyze coefficients in probability scores and coefficients in do we do coefficients yeah and coefficients of different words right so for example here for for business the words with the five highest coefficient of firm euros shares bank makes sense right and then for entertainment it would be film singer TV etc so this way they can actually like see an example put it into perspective what like the abstract things that I talked about before like the schools and and well the word vectors yeah and that concludes that then the fine then yeah then I talk about word embeddings because you can't really teach NLP course and be like yeah we only did one hot encoded vectors however like I will like we don't use them like in the live coding session like we don't we don't use word embeddings we use the one hot encoded vectors that we talked about before but still I want people to be aware of them obviously because then after that after this because this is kind of like their base right and then after that if they want to go deeper and actually go and start exploring transformer models sequence to sequence models generative AI, et cetera, then you can't really not know what a word embedding is. So I took a lot of materials from the book by Jurawski. So you will maybe see some familiar pictures, not sure. So how I introduce it is, well, in hot encoded vectors, right, all the vectors are independent from each other. They're all showing into different directions. So this is why we need word embeddings, and it was actually a student of mine who created this cute little graph that's actually we want something similar to this right so okay we want fridge to be completely like not complete but want to be kind of like unrelated and independent maybe from cat but we actually do want cat and kitten to show kind of in a similar direction right so this is how we would also then address one of the limitations of the bag of words model that we talked about during the very first session that now we can actually put things into relation whereas in bag of words you can't do that so then we talk about different vectorizers so this is where the inspiration is purely from Jurawski's book because I really liked it so we have the classical one hot encoded vectors that we already talked about so I'm not wasting too much time here going over that again then we talk about the term document matrix and there was a nice example in the book by using Shakespeare plays, Julius Caesar and Henry V, and that you can generate document vectors by checking how often certain words appear in a document, like in a book or in a play. So, for example, As You Like It and Twelfth Night, they have the word fool very often, but the word battle they don't. And so when you have bigger documents, exactly the same just that we can't plot it behind three dimensions right but this is basically how it works this is how you can find similar documents just by comparing basically how often certain words co-occur with each other or like which words occur and which ones don't and then you have this way you can also kind of already like create like word similarity measures right so we can kind of see that the vectors uh full and width are more closely related to each like they look more similar right based on the numbers than for example like good and battle but obviously not really a great way of doing it given that you know if you have like a really big book and a really small book obviously like the numbers will differ even though they might be you know from a similar genre so another way then this is where we already get closer you know in order to like you know spoon feed them later how for example um how skip gram model works is you know by talking about co-occurrences of words like the term term matrix right and we can see for example that the word digital and computer occur more often together than for example the um the word digital and sugar right so this way again we already like introduce the the the definition of like a context window right so to use a context of for example plus minus four words and more most likely you know similar words will co-occur together yeah so that's that and then I start talking about word embeddings so again what is a one hot encoded word we have for example a one for queen right so this word like this this is the representation representation one hot representation of the word queen but what do we actually want in order to capture similarity so we want something similar to this right and again obviously this is not how computer works it doesn't like tell us oh okay the first um the first number stands for royalty the second for masculinity etc but this is again i mean this is more or less uh how how word embeddings work and so i um so for example we for royalty uh king and queen and princess will have higher values but for woman not because it's woman is not royal masculinity king and queen sorry king will have a high score uh woman uh princess and queen will not etc but so this way we can already capture similarity right so this is what we want so basically we want to get from here to here and that's the difference between you know just a one hot encoded vector and an actual word embedding and then okay the famous example right which actually only works in perfect conditions we all know that but but theoretically you should be able to do vector maths right if you have if you have word embeddings right because you could for example take the the vector for king at the vector for woman and you will get to queen right um again famous example from from word to vec paper obviously it does not always work like that but theoretically this is what we would like to be able to do then i very quickly talk about statistic static versus contextual word embeddings so obviously contextual word embeddings are not like not within the scope of this of this little workshop so i'm not talking about again transformer models uh bird etc but i still like again mention it so people then can go and make do their own research or you know start like do maybe another course a more advanced one but i do talk about watervek because the algorithm is you know quite straightforward and simple to follow and we already like gave all the prerequisites for that right so we talked very quickly or like we talked about an mlp where we show that it has a hidden layer that kind of like loads the weights um we introduced a context window that um like the shifting context that can capture context and so yeah so i talk about word2vec, the two algorithms, the skipgram and continuous bag of words, right, that you either feed in the context and you want to predict the, well, the word in the middle, or you give the word and you want to predict the context. And using, you know, a neural network architecture for that, you can, you know, you learn the weights, right, and these weights in the hidden layer become your word embeddings and then because again I mean I do still want them to use them and not just again because I don't like stuff that is like super theoretical but people can't actually apply it then we actually also like train our own word embeddings but we are not using them to then feed them as I can to you know into a more complicated architecture we just use them to show cosign like to use cosign similarity because again it's simple it's easy to follow and and again, you know, and you can play around with it. And so, for example, one example that I give, imagine you have a chatbot for IT support. And then, you know, somebody types something in, and then you could use cosine similarity, maybe, to then extract something from the database that somebody else already asked, which is very similar. Like, I don't know, my internet doesn't work. But then somebody else says, I don't know, my router doesn't connect. Something like that. And that these sentences will be more similar than something that has nothing to do with the internet so yeah this is where we use gen sim and we use again the BBC data set in order to then train word embeddings right then in gen sim you have something nice which is called phrases so if something actually appears very very often together you can actually create like a feature you can create a token with those phrases and then replace the co-occurrences with those phrases and then we can train a model so using I deleted stop words just to yeah for for demo purposes so that doesn't take that long the parameters of the like of the where is it yeah how to how to how to build a word2vec model so the minimum count right which is the word code currents in the corpus then the the window what we talked about earlier plus minus four plus minus three again you can uh experiment with that the vector size how long you actually want your vectors to be and how many cores you want to use while doing that and then yeah so this way you can train your own um on word embeddings and so this is how they look right um so very very different from what we're used to from the first couple of lectures right so for example here that's the word vector for football and i talk about that you know every time you run the notebook they will be different right because the weights in the model are initialized randomly so and then also again logically from that follows that you trained your own word embedding now using your corpus so it will only work for you and this corpus right you can't then take like some downloaded ones and be and then wonder oh why is the vector for football so different um from from mine vector for football so this is really nice for stuff where you have like your own copra right so for example again in uh if you have like if you're dealing with legal text maybe worth training your own word embeddings instead of downloading something that was trained on google news or something like that um so this is where we talk about that that depending on the topic and the problem at hand you might either want to um yeah train your own word embeddings or use pre-trained ones and so they can also download the pre-trained ones i'm not going to do that now um it's uh so it's the word to like google news um word embeddings right they were of a size 300 and then again like i showed them that this vector will obviously look very very different because it was trained on a very different data set much more data than and well again weights are randomly initialized and that you that you can only do then um comparisons of similarity you know using obviously the same word vectors from the same corpus okay that's that kind of close it thank you and then we talk about your applications right so for example you can use cosine similarity um in order to calculate the um the well the cosine between two vectors and this way you can tell as we saw an example with the cat and the kitten right how um how close like how similar the direction is into which they point and so here for example we can see that the between the the the value between ball and match is much higher than for example between match and president and then obviously the question comes up okay but how do you do it between documents and so here this is where I talk about the fact that okay in order to again use using it is doing it naive way because the non-naive way is not is out of the scope of this little course the naive way is just to to average them out but obviously this has also its limitations right because once you once you did that you cannot get back to the individual vectors so you could actually have like some very different numbers here but you would generate the same like if you would average it out it would generate the same number but the vectors were could be very different however in the models that we cover here you know like that you use in scikit-learn there is no way to use like to to feed in something sequentially so you need to somehow like we did with the min sorry with the max and with the sum pooling with the linear sorry with the 100 encoded vectors we need to do something similar with the word embeddings um however i mean it again it depends on your task it might actually work quite well um yeah so we create the way we create the weighted sum and this way we can actually then you know calculate the weighted sum of the document word embedding and we can i used like the weighted sum of three documents what were the three documents i think yeah business article and two sports articles and then again i show them that the two sport articles will have high similarity than when you compare the sports article with the business article and now they know how to use word embeddings and yeah then a comparison between yeah well you could argue yeah well okay but we could still um use a cosine similarity for 100 encoded vectors of course you can right but as an example we have two sentences since i enjoy spending time with animals i decided to buy a dog and we have you are thinking of adopting a cat because you like pets, right? So you see there's absolutely no overlap between those two sentences. So it will give you a cosine similarity of zero if you use one hot encoded vectors. However, if you use the word embeddings, you actually get a similarity of 0.7. And here we get zero, right? Because there are just no overlapping words. Yeah, and this kind of concludes it. And this way they have like a nice little, well base for then continuing their their NLP journey and then as I said what I do usually is and this is actually we can do that together now is that I do a live coding session like from scratch I mean I do have like something on my other screen right so that I don't waste too much time you know with typos and stuff but and I also try to change it up each year either I obviously have some backups so do your homework if you're actually going to be teaching you know have some interesting data sets because you know if people are very new to this they also sometimes struggle just to find something cool well actually stuff works and so but usually then again I also ask whether somebody has just like whether somebody found some interesting data says something they want to cover specifically sometimes have people that have ideas because of work because they're working professionals and they say oh I'm dealing with this kind of data can we maybe do something with that over something similar which is obviously like not proprietary and available online and and yeah and then we just apply the stuff that we have that we have learned so very simple little data set from from Kaggle it's disaster tweets I think it's even like considered you know like natural language was like my first project or something like that right so getting started with competitions so because again I tell them it doesn't matter whether then you somehow get like you have bad results in your baseline and you get amazing results. All I want you to do is get your hands dirty so you know how to do it and also to have a template, right, because that's the beauty of this. Once you have a good template, then you can, of course, just always open it on another screen and just copy-paste or, like, follow the same steps that you did before. So I downloaded the data set, and then first thing, so, right, first just sanity checking. What is missing? Do we have missing data? I usually use the Seabourn library. okay location we have apparently like we have quite a few columns we have an ID we have keywords where we actually don't know how they were extracted location text and the target and we can see well location is missing probably not worth using that column then we really again check whether we really have two classes as we expect disaster non disaster because again maybe there's some you know some mistake or your data is I don't know got corrupted so just sanity checking. Then we check for the distributions to see, you know, are we dealing with more or less a balanced data set because that can obviously also, like, then that affects model choice and also, like, baseline choice. You know, if it's, like, completely unbalanced, it's kind of maybe don't use naive base as a baseline because then, yeah, you'll get, like, a crappy baseline and then everything else will be perfect. Then maybe use a different model that isn't as affected by imbalanced data sets then we set because it's a Kaggle challenge then we we already have a test set right we don't have targets for it so the way we set our baseline is we use we use a okay live base in this case we without doing anything to the text without doing any pre-processing and nothing we apply it and then we upload it to Kaggle and we see and this is our baseline in this case if it wasn't a Kaggle competition then as we learned it in the previous notebooks then you have like you decide whether you have a train and a validation set or you just use cross validation you take test set aside and you really just apply it once and then you put it away until until the end right and make base all your further decisions on the validation set or on the or using cross validation so you uploaded to Kaggle and a Kaggle competition baseline with a multinomial knife base without doing anything to the data is around 80%. So now, for all the other decisions, we need to use cross-validation or set a validation set aside, right? And now we can start experimenting with the parameters of the vectorizer, for example, right, or with pre-processing. And so here, this is usually where I spend some time, you know, like going through different values. For example, should it be true or false? So that's, you know, whether we want to use max pooling or sum pooling. How many times do we want a word to appear in the data set? So usually it should be set to minimum two anyway. Do we want to delete punctuation or not, right? Do we want to include a token pattern that we include, you know, everything and don't let the scikit-learn count vectorizer delete anything, right? And then, you know, you rerun this code with different combinations. I mean, you can do it manually. You can, you know, do a grid search. I'm doing it manually just to show people that the number here changes. And this is what I mean with, you know, like, setting a baseline correctly. That, like, don't just blindly apply some preprocessing because it will be different for every data set. Then you can also experiment with different vectorizers. So these are all parameters in your overall model, right, that people, like, sometimes just ignore for some reason. And then just really focus, like, on feature engineering, whatever, but you already actually didn't check whether, yeah, as I said, maybe including stop words is a better decision than deleting it. Cool. Well, then we, again, importing different classifiers, seeing whether that changes anything. So usually, I mean, my go-to ones are support vector machine, KNN, a decision tree, and an MLP. You again try them out and see, you know, whether any of them gives you a better result than before. Then, obviously, I try to choose one which actually has a few hyperparameters to tune. So that's usually how I also have my backup data sets where K and N hyperparameter tuning is really boring because you only have the K. But I usually try to choose a data set where one of the more, let's call it sophisticated models, perform a bit better to actually talk about hyperparameters as well and how to tune them. So usually I go for logistic regression or support vector machine. MLP I also try to tend to avoid because it's quite boring just reiterating all the different hidden layer sizes and neuron sizes, and it can take a really, really long time, and you don't really want to stand there for 20, 30 minutes, and depending on the size of the data set. So usually, just from my personal experience, choose logistic regression or support vector machine. Cool. Then what else can we do? We can actually, again, get our hands dirty and analyze data also qualitatively, right? So instead of just focusing, oh, this model did this, and oh, I'm getting hooked up on hyperparameters. Maybe also have a look at your data to see actually what went wrong. Because again, you will have, I've noticed, I was supporting a student a couple of months ago with her master thesis, and she found a data set in Hindi for fake news detection. And we couldn't really get our accuracy much higher for the model that we used. We also used, you know, different, like, more sophisticated models. We used transformers as well. And then when you actually i don't speak i don't speak hindi but she did and but when you actually then go into your data there was just mislabeled data like a lot um so again it's not always your fault or the algorithm's fault i mean also please qualitatively engage with your data which also quite often people don't do when you read when you read research papers right because again maybe maybe it wasn't actually the model maybe you had just bad labelers and maybe also didn't follow a proper labeling paradigm right who actually was labeling the data did you actually um like did you did you analyze how people like i've got the word for it um inter annotator agreement did you actually measure that because quite often they don't they just like either i mean if they um they quite often it's just the researchers themselves who do it and they don't even like especially if you like um if you if you don't have that much funding and can't really outsource it to Amazon Mechanical Turk or those other platforms and then it's like okay but who actually verified how good your labels are so yeah saving wrong predictions into a file for analysis confusion matrix all the good stuff then we talk about hyper parameters and so again for because we chose a sport vector machine we can search for the for a more optimal c and try out different kernels so this is also where I introduce regularization what is regularization why you should use it especially if you have badly labeled data you know then maybe you actually shouldn't should use more regularization and not trust your data so much so this is where I talk about that and then you try it on a test set and you upload it back to Kaggle and then obviously depending on how much time you have you can like go nuts and for example use what I also like doing is especially if the if the group is a bit stronger and have like more coding experience you can use spacey for named entity recognition, and then maybe introduce additional features like, ooh, it has a location, right? Or it has, like, some other named entities and generate features like that. What else? What I also like doing, but I don't have it as a notebook, but I have it like as additional code, again, if the group is stronger, or if I use it for, or if I use it with private clients, I also feed in the same data into a more sophisticated model, and then we compare and what is quite fun is when actually the sophisticated model isn't much better and then you're kind of like oh look deep learning is not a magic pill haha yeah so that's that so this is usually really like up to you how you whether you just do it like I did now and spoon feed them with a with a data set or whether you actually you know ask people to provide you with something that's that and then I talk about some other additional stuff which people then can either look in about sorry so first additional resources so again I created those notebooks as I said in order to address certain shortcomings that I've noticed in other online courses but I will not create content from scratch which is already really good somewhere else so there is so again that's literally just for you if you if you have ever actually either gonna like learn it yourself or teach it uh some cool links uh to to uh let me see what i find the one that i wanted to show you i really like this guy maybe you've heard of him it's stat quest it's i can't remember which uni he is he's a lecturer at some uni his videos and his explanations are amazing there's no point me wasting time if I can just show like this video like reference this video so this guy like for basic ML stuff it's it's it's a goldmine so yeah so some links to them but I mean I only included the links for the stuff that we've covered like knife base logistic regression but then obviously you can then you know just in general just check him out and so this is how I just provide my students some additional resources so they don't go and pay some weird data science course if there's like yeah no a lot of a lot of materials good content already online so pca step by step you know see um cross validation and like his animations are just spot on so yeah big fan uh don't waste your time uh you know explaining something or creating content if you can just if you've already found something good online um some links to courses so i really like you know just again you will not become a psychic get learn expert just because you know you watch a video but sometimes it's just good you know to have like a longer one for four hours that also runs you similar to what I did here through know a whole pipeline right like how you do the pre-processing how you do know it like normalization scaling if we're not talking about NLP then feeding it into the model etc so basically just you know this you can adjust if you have some amazing content please put it out and please share it with your students then proposal so each year I collect I collect like ideas or just you know wishes of what people would like to cover on an online course so what I really like topic modeling using top to back because again just from my personal experience it works quite well I used it for for for a project and then paper to it and then what I've and I get this again like as I already said previously it is a very common problem that people want to address is named entity recognition and in spacey doing them every recognition is also beautiful that you like easy and so I show here what that means right that you have some text and that you know by using spacey you can get quite a few entities out of it but what if you want so here you can see a visualization of that but what if you want your own ones right as i said what if i wanted so for example in in my legal tech job we were interested in certain drugs i mean that's a bad example there were like more but we had like hemopathic you know certain vaccines certain other stuff and um but um the dictionary was I mean we also wanted it to generalize more right so we didn't want to just use you know if check if word is in but we wanted it to generalize so we trained a custom named entity recognition classifier so again and in spacey space is also beautifully documented it's just so nice to use so this is also something that I could cover but I would not need two days for down right so it's two three hour thing um so yeah so i also for you guys if you're interested there's a form here that you can fill out then there is my my my linkedin my websites and um yeah and that's kind of kind of concludes it we still have quite a bit of time so if you either want to go over something in more detail or if you just want to have like a longer q a session or just go and grab some food that's up to you but I will give now the word to you guys thank you very much for for hanging in here for an hour and yeah Because I don't see any in the internet, so there's none of it. So can I ask a question? Yeah, of course. So for some models like that... Hm? Yeah, sorry. We have some models like that. yeah sorry we have some models like that do we need to remove the support in the train data for bird no because bird is a contextual model so it actually you need to keep the stop words but what you can for example do is again if you have typos in there or rare words then you could for example remove those right but as to give you an example like a certain hashtag or something in twitter data but no in bird um you do like very minimal or like in other transformer models you do a very minimal pre-processing that's a good question i would do it for normalization yeah um because again like you just because again at the at the end of the day bird is not a human being we don't need to make sure that bird understands that something is a is a is a url so actually doing like normalization you know like saying that oh this is an http address so that can you know learn that feature i would do that so if we can't find an event or a link we need to remove them and put it a kind of label exactly so that's what i did here but again um so this is why i was like kind of like you know religiously saying throughout the talk always validate you try it with and you try it without and see what works better right i'm not telling you now oh this is how you need to do it right because it depends on the data set so set a baseline without touching anything and then try different pre-processing techniques and see what that performs better yeah that's that's um yeah can you maybe tell us a little bit about some common difficulties you have when teaching people like these nlp concepts like are there any topics that you find especially difficult to teach or so it really it depends on the setting so i can tell you so because i was teaching at university a lot and it's um i was teaching a data science so it was not nlp it was it was data science summer school a couple of years ago at ucl and we had a lot of people from asia and they um it's a bit like there's a cultural difference between you know certain you know different cultures um that they kind of they don't they don't they don't immediately understand that there's no one solution fits all approach you know like they see this and then they follow it step by step without you know actually trying trying out like different things and this is also what you know i tried to tell you to tell you guys earlier right there are different types of data use logical common sense um deciding on yeah whether you want to remove punctuation for example or not whether it makes sense but they kind of they want like a very strict you know first you do this then you do this then you do that and you kind of need to like you need to explain to them no i don't know like you know you need to experiment and try things out there is no one solution fits all approach so i think you know but it's it means it's like that for every topic like it just comes with experience at the beginning it's obviously easier just to follow a certain certain number of steps and instead of you know like using your experience and common sense to make decisions as one example I will give you I'll give you some examples as I said from my from from a previous project I worked on so there we actually used a lot of linguistics to get information out to give you an example I mean grammar is actually quite logical and structural so we were trying to get out which part of a regulation applies to which kind of drug this is why we were trying to you know identify first two drugs because EU regulations or UK regulations not like you have a regulation for each type of drug you have like the Medical Act of whatever 1994 right and it talks about a lot of different medicines then first you use a named entity classifier to find the paragraphs that talk about certain drugs and then we use parse trees in order to find out whether a paragraph applies or doesn't apply to a product so for example this paragraph applies to homopathic drugs comma but doesn't apply to and then by you know like building those like your your grammar rules you can extract information out like that so quite manual but it works And then at the end, the use case is that a user is interested to figure out how to address the user. Yeah, exactly. So, for example, a medical app comes up with a new drug, but we are too poor to hire a lawyer or we just want to get like a first impression of, okay, which paragraphs do we need to be aware of, which regulations do we need to follow, instead of, you know, like reading the medical act from start to finish. So that was one thing that we used. Another thing we used, so this is why I had like those two examples up there. we were working with DEFRA, which is a department, like an environmental department in the UK. And they already had some regulations that they said, oh, these regulations, they have something to do with climate. So they, you know, they address water safety or air pollution. Can you find more? And then use topic modeling. And you check those regulations and which topic clusters they are, and then check the documents around them. Because most likely they will be like quite similar in this way you find more regulations that apply to yeah air pollution water safety etc so these would like two things that we were working on I mean regarding interface that was that was not my thing that was outsourced to the Eastern Europeans who they made a pretty interface no I'm kidding and so what you're talking about like obviously another big thing in law is actually finding out whether for example as what exactly is an argument what exactly is the justification what exactly is an example you know why certain decision was made right and there I mean it's you need a lot of label data for them right and then like it's you can't really go really far you know using you know like a manual approach like that following a path tree so we want for example like another startup that I'm sometimes supporting they actually want to you have labeled data for which trial was successful and which wasn't, right? And then you just need labeled data and, you know, really powerful contextual models that can distinguish, you know, between, like, ooh, this trial and this trial, what are the similarities? You can't just, you know, use an approach like, oh, this paragraph applies to, right? You have actually, like, one is a, I don't know, One was a homicide, and one was a murder. Two different decisions were reached. Why? And then you have a third case, and then you, well, which one is the one that you should apply, right? Should this be a successful outcome or not? It's probably very hard to validate, too, right? Yeah. So that's just law-specific. That's a big thing that people are working on right now. I have a question about chatbots, because I'm also trying to make chatbots sometimes. What is the basics of making a chatbot, like a simple one, not as sophisticated as GPT? Okay. How to tell in terms of the methods that we, that you show us that we've seen, like, how to introduce people into... Into chatbots? Into chatbots. That's a good question, because I did my PhD on that. So it's a very simple one, right, there are two ways, and both of them are actually covered, right? Either you have a chatbot and you have some sort of data. How you get it is up to you. I actually recruited people because there was no data set available. So I had to collect basically the knowledge base of the chatbot, like potential questions that the user might have and potential answers. And then there are two ways of doing that. Then one way is using cosine similarity, saying, okay, if this... Well, actually, the more simple way is just if-else statement. You connect them all. It's like, oh, if this question comes in, or a similar one, then reply with this yeah um or again you use cosine similarity okay something this question is similar to another one i know what that we need to answer with that one you measure the similarity then um that's one way the second way is you could train a classifier right and so to give you an example my last my very last paper was on a chatbot that tried to convince people to get the covid vaccine and yeah exactly how to do that and um because there was not a lot and still like there is not a lot that much information on like no people are not that informed people have like very specific reasons why they don't want to take the vaccine so it and they weren't super super sophisticated and well research was just like um it was developed too fast it will affect my my fertility that was a common one with with women so it was always they used kind of the same keywords or somehow like the same ways of expressing it and then you could just label it right and say issue fertility issue chipping issue um issue uh to develop too fast right and we we call this concerns right concerns that people have and you could train a classifier that classifies the concern and then tries to reply with a counter argument to address your concern That means once it sees the keyword chipping, it starts talking about it. Yeah, exactly. But keywords would be very similar. It would be just like a control F, right? But sometimes people would use different words. So for example, they could use side effects, but then they could use... What was another word they used? There was another one. So the words would be different. And then obviously if you have a classifier that you train the data. 5G, for example. Yeah, exactly. Exactly. And then for deployment, I was thinking also once making a tutorial about because actually it's very easy to deploy in telegram. I really like it. I mean, I didn't because it's, I mean, there's still like, then you have the boundary of people needing to have telegram. So it's easier to start a website on the server. But on telegram, it's actually really nice, a really good API. And then you can can host it there just for fun. You just do something like digital ocean as a server, right? And then it just, yeah, it's very neat. Not sure what they've done it before. They have like the bot for it. father, right? And then you create your, yeah, it's really neat. I have to admit that. Yeah? That's one idea. That's what I was, I once did an interview, and the guy came up with that, and I was like, oh, that's actually really cool. That's a good idea. Okay, so if I have a medical data, so it's sensitive, which kind of method do you use? That's a good question. I mean, I would probably just buy an API, right? I mean, don't post it into Google Translate. I think DeepL has an API. I'm not sure whether they're storing it. You would need to literally just read their terms and conditions, right? Whether you yeah whether you should be using their I'm not sure either to be honest I have not dealt a lot of translation yeah I mean probably you could train your own translator not sure I would do that and go through the hustle I'm sure they're like translators you could use with decent terms and conditions that will not store your medical or you could anonymize it right so in medical data that's also like a big topic that you're actually even for your models you don't want to have any names and stuff in there or any like personal identifiable information so it's another that's another area of research how to actually you know anonymize data properly so it's not just about finding mr and mrs right but if you actually for example know the location it's a small it's a very small village and you say it was somebody of again certain descent if there's only one person of that descent living in that village well then it's kind of obvious who it is right so cool okay guys I mean we still have ten minutes I'll be hanging around here so if you have any questions that you don't want to be recorded find me so yeah but still thank you very much for coming to my talk I really appreciate it

Dr. Lisa Andreevna Chalaguine

About — in the speaker's own words

Originally from Belarus, I currently live in London and work as a data scientist at ProcureAI. I did my PhD at University College London, where I used to develop algorithms for chatbots that can engage in argumentative dialogues, trying to persuade the user to accept the chatbot's stance. My expertise therefore mainly lies in knowledge acquisition/representation and, of course, natural language processing. I also have over 10 years of experience in teaching and tutoring. I am also a tutor at The Profs, and a member of the part-time tutor panel at the University of Oxford in the Department of Continuing Education.

Social card for talk: How to teach NLP to a newbie & get them started on their first project