Episode 81

How To Break Into Designing "Inclusive" Microsoft LLMs + Navigating Bias & Privacy Concerns - w/ Advitya

Oct 16, 202401:09:41Audio-first conversation
How To Break Into Designing "Inclusive" Microsoft LLMs + Navigating Bias & Privacy Concerns - w/ Advitya cover art

Advitya works on Responsible AI at Microsoft as a Machine Learning Engineer 2. I am extremely stoked to present this incredibly nuanced discussion on the ethics and responsibilities that come with building mass-market generative AI tools.

Who this is for

  • You are building something real and need the unpolished version before the launch post.
  • You would rather hear Advitya's version while the mess is still fresh than get another polished hindsight sermon.

Key takeaways

  • Break Into Designing "Inclusive" Microsoft LLMs + Navigating Bias & Privacy Concerns - w/ Advitya
  • How do we make sure that whatever the model outputs actually grounded in real-world data and is actually correct? How...

Fast scan timestamps

00:0000) Intro + Background(
00:0351) Advitya’s career snapshot(
00:0522) Landing Microsoft’s ML rotation program role spanning 4 orgs(
00:0825) Experience of studying Data Science at UC San Diego(
00:1111) Day in the life starting out at Microsoft(
00:1452) Hands-on experience working on products at Microsoft(

Transcript

The full conversation, right here. Auto-captions, lightly cleaned, still very much a real human conversation.

11,249 transcript words114 transcript blocks
Speaker

You can think about responsible AI as a big umbrella of safety guardrails in that you want to bake into a model across different pillars like bias, transparency, privacy, robustness, interpretability, etc. If you were to pull a sample of 100 people, maybe 60 would say that, yeah, that's problematic, 40 would be like, yeah, that's okay. Who decides whether that response is in fact problematic? Is it some shadowy leader that you never meet or is it you as an engineer or like, who is the decision maker here?

Speaker

I got to work with machine learning systems research to AI for accelerating IT operations to even responsible AI in my final rotation. Maybe I should have led with this, but where does responsible AI come into what you just described? How do we make sure that whatever the model outputs actually grounded in real-world data and is actually correct? How far do you think we are from Agi? Welcome to the Ready Set Do podcast, where we learn from journeys of not experts who are just two steps ahead of us.

Speaker

I'm Naman Pandey and in this episode featured not expert is Advaitya Gamavan. Advaitya has the incredibly unique role of working on responsible AI at Microsoft as a machine learning engineer to more so than the other episodes that I've put out here. I'm incredibly stoked to present what I feel was such a nuanced discussion on the ethics and responsibilities that come with making mass market generative AI products. This is a discussion which is oftentimes super technical and complex, but I believe one that can really change the way that you view and approach generative AI.

Speaker

We discussed the technical foundations of Advaitya's career that brought him to this cutting edge of technology where he now lives in dreams and also go over how Gen AI tools by big tech navigate critical concerns such as privacy, ethics inherent biases that creep into data all the time apparently. And the aspect that I was most curious about, who are these people that are the actual decision makers when it comes to flagging a response is something that's problematic and or offense.

Speaker

In keeping with our theme of learning from somebody that's just two steps ahead of us, sir, I'm an expert. My goal is to highlight the extraordinary role that Advaitya is playing as an ML engineer by working on something that is the harbinger of the next era of technology. Indeed, the future is now old man. This is the ready set to podcast and to support it. Please subscribe to the YouTube channel and or leave me up to a five star reading spotify and now without any further to do my friends.

Speaker

Here's a welcome up with you for having me in a month. It's great to be on your podcast. I think I came across your videos a couple months back. And honestly, it was just a lot of fun to watch. I kind of love the amount of detail that you covered in the podcast and like things about both the professional and personal journeys of people, which kind of like made their overall journeys very authentic to me personally, as I got to learn about like people in different fields.

Speaker

So, yeah, excited to like chat with you today and I hope this conversation also brings in comparable value to your audience. So, so a simple metric that at least I use. Maybe it's not the best one out there is that I just optimized for stuff that I want to know really, you know, so with you, I'm really excited to get into what the fascinating world of responsible AI because clearly that's the direction that we're headed, but going back to what you said, it's not something I've heard somebody give that feedback to me which in which you reference that it's both the personal as well as professional side of a guest that comes out.

Speaker

That's not even necessarily something that I tried to do but so that's what makes that feedback really valuable to me, because now I know to, you know, explicitly try and do that more often. So really appreciate that call out and yes, just to get the ball rolling, do you want to give us a, you know, quick and dirty snapshot of your career that kind of got you to this point at Microsoft working with responsible AI. So my name is a good team of it. I am currently working as a ML engineer at Microsoft. Currently, I build tools for evaluating large language models background I graduated from UC San Diego majoring in data science.

Speaker

And then I have been working at Microsoft and it's been a little over three years now. For the first two years, I was in a rotational program. So it was actually a two year rotation program where I worked with four different orics across Microsoft for six years each on different kinds of ML projects. I got done with that program summer of last year. And since then I have been in my current team working on different kinds of responsible AI and AI safety tooling.

Speaker

Initially, I worked on tooling for computer vision models, and this year onwards it's quite a lot about large language models. So yeah, that's my intro in a nutshell. Yeah, appreciate that love how, you know, actually quick and dirty that was, you know, it's, it's a very delicate balance to strike, which I think you did really well there. I am curious, what were the four divisions that you had to rotate at? That's the first question. And the second was that that actually sounds so cool to me.

Speaker

This is not something I've heard that even happens at big tech. So I'm also curious how you kind of landed or found that position if you could touch on both of those. To answer the first part of your question, I got to work over the span of that rotation program in four different orics. I personally chose to work in four different orits across Azure, because I really wanted to be involved in like different kinds of product lines within Azure and solve different kinds of machines.

Speaker

Okay, so I got to work with Azure quality, the action platform aspect of Oregon, but I also got to work with Azure data, which kind of like relates to building their data products and product line. I got to work with Azure, which actually is related, which actually is related related to building the actual ML platform, which companies can use to actually build ML models and productize them and put them to production.

Speaker

So I landed this program. So, I actually had done research with a professor for two and a half years in college, and they actually had very strong ties with a research lab at Microsoft. When it actually, when I was actually searching for jobs, right, I kind of came across this rotation program, and I reached out to the head of that lab just on LinkedIn, introducing myself and saying that I've been working with this professor.

Speaker

And like I know that you have very close academic ties with this research lab I've been working with them, and I would love to perhaps get a referral or like hop on to an information call. And if they, not only did they know the program, they were actually working with that program, actually. It's so long had like one or two kind of projects every semester with that with the program that I then got to be a part of. So not only did they give me a referral, they actually shared my resume internally with the hiring manager of that program who went on to be my manager when I joined Microsoft.

Speaker

So, well, that kind of experience felt like even more enhanced than a referral where someone on the inside was actually vouching for me as a potential person that should be interviewed. And that is how I entered the interviewing process. So, yeah, it was pretty groundbreaking and to also kind of elaborate on how I came across this program and things like that. Of course, sure I was using typical tools like Google and LinkedIn job alerts.

Speaker

But what particularly stood out to me about this program the month is that I couldn't really find any other program of this kind in the industry. And just for some more context, what did you study at University of San Diego? I actually majored in data science. So, okay, data science. Yeah. So data science was a new major that got introduced in my college when I was a freshman. So that is when I kind of got to explore more of it and I got to be one of the first cohorts of that major at UC San Diego.

Speaker

And did you have a lot of hands on experience working with machine learning during your college days or was more of that happening once after you'd already started your stint at Microsoft. I do feel pretty grateful that not only was the major view, but a lot of the classes were really tailored towards data science and machine learning concepts. A lot of the required coursework that we would have to take was a combination of math classes, the typical CS classes and pretty much machine learning classes. We had dedicated classes to deal with different parts of the machine learning life cycle, all from like theoretical foundations of data science to practical applications, using some basic modules and frameworks to build different machine learning models.

Speaker

I got to take dedicated classes in data engineering, data visualization, analytic systems, cloud technologies, things like that. So recommend your systems, things like that. So I really kind of like loved how tailored and specific a lot of the classes were in terms of dealing with different concepts. That is one part of it of my experience. The other part was that as I look back at my experience, I feel it gave me a pretty solid foundation in terms of how I practice nurse actually build AI in the industry.

Speaker

But I generally would like to emphasize on the word foundation, because things are moving so fast in our industry that I would even like have to be honest with you and say that pretty much all that I learned in my college was very useful foundational knowledge that is now deep traditional in our industry. Wow. That's so impactful. Yeah, I mean, that that makes so much sense. Yeah. And then, so from there. Yeah, this is so interesting because what I'm envisioning is, yeah, you, you know, go to college, you do these really hard courses.

Speaker

They sound very hard to me, even though I do have some background in computer science, I was never as deep into data science and all of those things like so a lot of that is still a black box to me, even though it's really not because I've actually had some exposure to these technologies at least to a decent amount at least, but what I'm getting at is, once you started your, you know, rotational program, and then, and say, I think your reference towards the end of that, or in the second year you were exposed to some of the more, you know, machine learning oriented stuff at Microsoft.

Speaker

So what was a day in the life of for that like, what was it like for you to hands on be working at one of the biggest companies in the world and contributing to their artificial intelligence research, or maybe it wasn't research, but yeah, any thoughts on that. A disclaimer, my experience is going to be likely not the norm, but an exception in the sense that, because I got to be a part of that rotation program, my experience was likely quite different from a typical you had experience, not just necessarily in Microsoft, but typical larger tech companies and all for that matter. So, in my rotation program, we basically had four rotations of six months each. So we were doing six months long projects.

Speaker

And the dynamic of working there was quite distinctive and interesting at the same time in my personal experience, where we were in kind of like teams of say, one PM to data scientists to engineers from our rotation program. Everyone's a new crowd at varying levels of degrees. And we were kind of dealing with two stakeholders at the same time. One was a partner team that we might work with the partner for whom we are actually doing the project or kind of accelerating and over for them.

Speaker

And the second stakeholder would be our actual reporting managers, or the leadership team of that rotation program itself. So it was a pretty interesting balance of trying to kind of prioritize what the partner or it wants to do what makes the most sense from a business perspective. And at the same time, how to further our own kind of ambitions in learning, and also kind of making sure that we're kind of getting the, the kind of projects or work or challenges that we're also looking for.

Speaker

So that in terms of like finding our interests among the variety of different kinds of projects that were out there for us to choose from. And then also kind of making sure that we're also learning valuable skills in the process from like a career perspective as well, not just a technical perspective. In terms of what were, yeah sorry what were some of the things that we're actually working on and getting your hands dirty, but if you can talk about that course yeah so a lot of the technical projects.

Speaker

I got to work with different very like different kind of sub problems of machinery, all from research to AI for accelerating IT operations to even responsible AI in my final rotation. So I got to work on like both research oriented projects and like production kind of typical software oriented projects as well from a machine learning perspective. It was all kind of like diversified in all from like publishing papers or presenting at yep publishing papers and things to even doing like a product work or productizing like ML applications or typical ML solution.

Speaker

For specific domains of problems to even building different kinds of tooling and infrastructure for other ML practitioners to build better using those tools. I say, that's really helpful. Do you by chance have an example of a product it not necessarily that you worked on but just for me to contextualize that something research that you did around machine learning that then became a product, maybe something within the Microsoft suit that comes to mind.

Speaker

I got to work on a couple sort of research problems and engineering problems. Some of that stuff is already being integrated some of that stuff might be in like sort of like a pre release or beta phase, but probably the closest thing that I can tell you, which is like the closest to to release and being out there for external customers is built. Is in my final rotation, I got to work on responsible AI tooling for object detection models. And I got to work with the Azure machinery or it there.

Speaker

So back at that time, so this was early 2023, the major platform that was offered by Azure was Azure machine learning studio, which was easy this one stop shop platform to train your machine learning models. And, like test them and put them to production. Features in for related to a model asset in M O studio was related to responsible AI insights. And there was a responsible AI dashboard integrated there. That responsibility I dashboard is actually also open source on GitHub.

Speaker

So people can also download that locally and play with them and just plug their model and data in and actually get a whole bunch of responsible AI insights related to their model, and there are a whole bunch of models they're supporting. So I particularly got to work on option detection models, which for that we were trying to add in support for different kinds of responsible AI insights for those models. The kinds of insights and tooling that I got to work on and service were related to a adding support for metrics, different kinds of metrics, be adding support for group based metric calculations, where you serve as the empower to do their own analyses, based on their own domain in terms of choosing different cohorts of data, different categories within their dataset.

Speaker

And that kind of comparing and performance metrics across those different cohorts of data to do more granular analysis on where failure points are exactly occurring, and what kinds of steps or actions might be needed. For example, if the metrics related to object related to objects detected in images are in low light conditions, versus like in the morning, for example, it is possible that the overall performance among those different categories or cohorts of data could be different.

Speaker

Like more insight to the user in terms of what kinds of data augmentations do they want to do, or how much do they want to change their training or testing datasets to actually make their model more resilient towards kind of those kinds of see like images tested or those kinds of data points, that has bad lighting. Exactly. And this is one kind of tooling that I got to integrate. There were additional awesome. Yeah, there were additional insights related to say even adding support for saliency maps in the sense of if you take an image and the model.

Speaker

You could have see like colors super imposed on top of the images, to actually see how your model is behaving. What part of the image is the model focusing on more or less. When it actually comes to the model making its prediction, so that the user can make sure in terms of transparency and better interpretability of the behavior of their models on how their model is actually making the prediction. What part of the images and focusing on, and is that right or wrong. And there's these different kinds of like charts and visualizations and metrics that I got to kind of work on as well, which also then got released into public preview, which people can sign up and use on the platform as well.

Speaker

That's awesome. Yeah, I bet that felt like at must have felt like a milestone moment for you. It's pretty groundbreaking for me personally in a month because the previous rotations that I got to work on. Those were really kind of like early stage ideas and research heavy and new kinds of projects that we were introducing, which would gradually be getting rolled out internally before even reaching external customers. Right. But in this case, I kind of got to be the closest to production out of any of my other rotations in my personal experience directly working with external customers and talking to them and understanding their pain points and accordingly surfacing different kinds of insights and tooling support was a pretty groundbreaking experience in terms of my kind of tooling and product

Speaker

being used by other companies out there. Awesome. So about that really amazing deep dive that you just gave us. I think most of the technical details, definitely track by me. All of that makes sense. The part that I'm not 100% sure about is where it does or how exactly maybe I should have led with this. So that's probably on me. But where this responsible AI come into what you just described. So what you were saying about, you know, the user being able to decide what part of the image to focus on like the user can tell the model.

Speaker

Say we're looking at the picture of, I don't know, let's say like a mouse. Right. And then the user is able to point out that when you do your analysis, please focus on these, this little block here, which contains the mouse. So that sounds more to me like an image classification type problem. I'm trying to understand your perspective on where the responsible part comes into. This is detecting skis, okay, like a bunch of skis for people who are skiing. So it's, it can be a lot more likely for a model to detect skis, which are just flat laying on the ground, which may have people using those skis and which are around snowy conditions, for example.

Speaker

So it can be very easy to see like the background like snowy conditions or say even a person. And then if they have like a long shirt of, if there's like a long sort of like thing at the bottom of the person that's likely a ski. But also, if there are skis that are just kind of leaning against the wall, for example, the other external constraints like snowy conditions will be typically see skis or humans using them are not really shown in an image of standalone skis just leaning against the wall.

Speaker

And so in the sense that kind of data, that kind of augmented image can be to save lots of typical inaccuracies in terms of the model being able to detect it. Because the model did not focus on the on the on the qualities of a ski nearly as much as the environment where skis are more commonly located. In other words, the model, basically, the kinds of patterns that the model extracted in its training process wasn't related to the inherent qualities of an object.

Speaker

But more so the qualities of environments where that object is more commonly located. I see, that's, that's really helpful. And that ties the loop in terms of my understanding of, you know, your final project with the rotation. So since you graduated from that program quote unquote graduated. I know you continue to work on responsible AI still. So at any point, did you get to, you know, dip your toes, maybe, or maybe even just dive head first into the generative AI side of things as far as responsible AI is concerned.

Speaker

So, and just for some more context, what I'm getting at is we know that AI can inherently have a lot of biases and all of those things which it needs like somebody needs to go in and tell it that a maybe what you're learning from the data is having you have this bias, but maybe you shouldn't actually have this bias. So, for me, for, you know, just an uninitiated layperson, that sounds like it would fall under, or it should probably fall under the responsibility I bucket as well.

Speaker

But I'm curious to hear if that is actually the case, or is that a separate department. So kind of answering the question in two parts related to how I, how my kind of break sort of evolved, where LMS kind of came into picture and kind of the role of say bias, for example, in the overall kind of umbrella of developing responsible AI systems, right. So first, last summer, when I kind of got done with the program or sort of graduated and then joining the team, I was initially working on other computer vision models as well.

Speaker

So broadly, like typical models like for image classification and multi label classification and things like that as well. Since this year, I basically got looped into building responsible tools for large language models, considering the overall priorities of not just like a company necessarily but just the industry in terms of the kinds of AI and the kinds of use cases that they want to unlock with these powerful large language models right.

Speaker

So, since this year I have been working on a lot of tooling around that. And a lot of the tooling around that is of course centered on a variety of responsible AI pillars, one of which is bias, which makes a lot of sense as you correctly highlighted. So, the way that I would look at these different sort of concepts like bias or transparency responsibly etc is that think of responsible AI as a big umbrella term, for example, we just released building AI that is safe and reliable and transparent and doesn't encode human bias accurate accurate and does not harm does not unintentionally harm any users using it right within this broad umbrella of responsible AI.

Speaker

There's office, there's definitely a lot of pillars, a lot of sub pillars and branches in terms of how do we actualize this. One of the most as you correctly highlighted is bias. How do we make sure our model does not have bias towards any specific groups of data or more practically any specific groups of people across different kinds of categories like like race or gender or traffic, and things like that, right. So, that is also related to how do we make sure that across different parts of the training process but say, especially at the very initial stages, how can we sanitize our data in a way where our data ideally does not encode those kinds of biases. How do we track the performance on a category like bias through different metrics like demographic parity or kind of testing it on different kinds of bias detection data sets.

Speaker

We make sure that are that the outputs of a model, especially a large language model is sanitized at all stages of the process, all from the user sending a prompt to the model processing it to the model outputting something and perhaps additional post processing steps to make sure that no bias is really kind of outputted by the model like to the customer. That is one pillar of biases within the responsibility umbrella. There are a couple additional pillars as well.

Speaker

One of those pillars could be transparency, for example. So, in the context of LMS, transparency could mean, how do we make sure that whatever the model outputs is actually grounded in real merge data and is actually factually correct. An additional way of doing that and surfacing that for the user could be perhaps making sure that whatever the model outputs as citations around each sentence that in generating to make sure the user can cross check references and what the model is outputting and make sure things are correct. For example, yet another umbrella could be privacy. How do we make sure that the model ideally is not trained on sensitive data, or if sensitive data got injected into the process, how do we make sure it does not output that at any cost.

Speaker

Another variant of privacy is how do we make sure that the model does not leak its initial system prompts or instruction prompts or potential details that the model is not supposed to say. Well, and aside from even privacy or anonymization, making sure that the outputs are kind of sanitized for removing personally identifiable information, even if it might be dealing with some kind of customer data. So another bucket in the umbrella could be robustness. How do we actually test these models across these different pillars, for example, one of the most common kinds of stress testings or attacks against a model could be adversarial attacks,

Speaker

where there might be an adversarial prompt, which might want to instruct a model to do something bad. How do we make sure the model does not give into that prompt and does not do something bad, or decline to answer and things like that. And there's very interesting and creative base number in which people have tried to like jailbreak models, introducing some, oh, I bet that's like that during the way from its typical behavior. And so just to kind of summarize the answer, like within the, you can think about responsible AI as like this big umbrella of safety guardrails and practices that we want to bake into a cross different pillars, like bias, transparency, privacy, robustness, interpretability, etc.

Speaker

To really make sure that these models are safe when they go out. Amazing. I mean, just from what you just laid out, I actually have already so many questions in my head. I'm having a tough time, even, you know, having them pointed down so then I can ask you one by one. So just inspecting the various buckets that you mentioned. I think some of them are fairly cut and dry for me. And I would assume for most people such as, you know, transparency, for instance, if we pick at that one.

Speaker

It feels fairly obvious to me that, I don't know. Maybe you could, you could correct me if I'm wrong, but probably 99% of your average lay persons out there would probably agree that it is a good idea that any time an LLM said says something. There should be a citation that if should they won't, they can go and look at that. So I don't think there's a lot of debate on that. So I think the context, most of the time, I'm just having an interest there. If I want a model to just create a beach generate some sort of a poem, right.

Speaker

I guess having citations could be like a slightly weird user experience, or I just want a model to see something completely unique. In that sense, the model struggle that creative. Any on the right train of thought here. Absolutely. No, that's a great call out. Yeah, because if you're, if you're having a LLM generate, you know, like a picture of a cat under a night sky. Yeah, probably I'm okay, not looking at the citation for that.

Speaker

You know, I don't care where it got it from, but that's where we get into the privacy concern, which I know was a big topic recently, at least, you know, I think it was more so great with Gemini, but it really just applies to all LLM's. But I think, oh yeah, Google, I think was recently in the news for this where they've been using all available YouTube footage to train their, you know, a video models, and a lot of people had the opinion that that's inherently wrong, or that shouldn't happen.

Speaker

And then if we just even stick to the pictures part of it. Obviously, if you have, if you ask an LLM to generate, you know, me talking to you in the sky in the style of Van Gogh, that is probably not original art, because it's, you know, giving it's not giving you credit to an artist that should probably deserve credit for that. All of those, you know, sticky areas we get into. So I want to examine the privacy aspect of it first, and then maybe we can go into, you know, the bias, which is for me at least the biggest, you know, hot potato.

Speaker

So to speak in this entire arena. So what are your thoughts on in terms of what I just laid out where do you personally land on the whole privacy debate around LLM's. It is definitely like in the kind of scope in the sort of timeline of when LLM's came in, and their overall training process and their kinds of outputs. It's definitely a long standing and continued problem, continued challenge, right. So it's quite interesting where there's definitely say like data on the internet that was likely used to train these models, right. And it's possible that likely, you know, explicit permission may not have been taken for these kinds of models.

Speaker

I've seen I believe court cases where I believe certain large news media organizations actually sued some of these model vendors, some of these companies that create large language models for I'm not kidding. I think I've seen in the order of trillions of dollars worth of reparations cited in those initial court cases, saying that there is a whole bunch of copyrighted data that was likely to train these models and things.

Speaker

Yeah, so I think there are two things that are sort of broadly happening here. One, now moving forward, people seem to be much more aware of using copyrighted data and things safer, the training process. And so I have get many heard, say interviews of like senior executives at these companies that are, say, creating some of these models, saying that now they are trying to get explicit permission for training models. Yeah, and so they're making sure that the model is kind of copyrighted, or they're making sure, you know, scrub off all the personally identifiable information before training models.

Speaker

And even more, and even more effective approach here is that a lot of models are actually now using synthetic data sets to train their models, where they're actually going to say, initially starting at some form of high quality data, like high quality records, research reports, textbooks, et cetera, and actually using them to generate synthetic data sets to then train the model on to sort of like escape or bypass this problem of explicitly training on this on like copyrighted data.

Speaker

And so the other thing is that I believe I've also seen, and I might be wrong, but one of these large companies that kind of own like a lot of this video content and audio content and things, they're even starting to roll out programs where even if they have AI features mimicking, for example, the voice of say some of these like artists or content reviews, they're still kind of gradually offering them royalties, for example, gradually offering them royalties or kind of like partnering with some initial artists to actually still surface those kinds of features to users.

Speaker

And so, because of that, because of that I feel like now the industry is definitely starting to realize the impact of you know just like training on some data or like kind of being more aware of the sensitivity of like how sensitive like data could potentially be right. Which steps in the right direction. In terms of the stuff that's already been done right, that is something that I believe it might even go on to, of course, like the legal system and the political system to kind of also decide how to mitigate challenges like these, what kinds of guardrails need to be baked into the model for making sure that that data is not explicitly referenced in the models output responses, and what kinds of legislation should then get created in the process to kind of mitigate challenges

Speaker

like these, we've of course seen like places like the EU kind of driving some regulation in the space, there has been regulation in California, believe an executive orders some time back that have also tried to kind of make progress in the space. Yeah, you can always trust you to do any and all regulations, whether you'd like it or not. Yeah, but one sort of me more catchphrase in the month sometime back, saying that like us in a way you regulate China copies.

Speaker

There's nothing like good or bad about these rules in my personal opinion. But unfortunately, all these rules are kind of like uniquely important in their own terms in terms of actually democratizing technology making it impactful. I remember I wrote some time back in an article that I would say India scales because of its unique kind of qualities. But yeah, it's pretty interesting to see the different roles of these different nature players.

Speaker

It's very true. And maybe this is just my own subjective opinion but for whatever reason I, at least in my head I find it difficult to separate the role of regulation and just slamming the brakes, you know, that's just what the first vision that comes to my mind when somebody says regulations, maybe that's on me. But coming back something that you said really caught my attention was you said that to generate that synthetic data.

Speaker

There's documents or textbooks that are being used and I'm assuming that obviously permission must have been taken from the right for owners of those things to generate that synthetic data. But what about pictures though, because don't you run into a fresh set of problems where like how do you get sensitive data on pictures because even if it is synthetic, it must come from something right and if it is, then we're back to the snake almost eating these old days. So it's my understanding right there or is there a gap there that I'm not aware absolutely on the right track.

Speaker

These are like I'm not going to claim to know the answers to all these questions but these are definitely quite important and complicated questions to answer. There are the most clear answer that I can personally give you the mind is that there are certain tricks that can kind of do any use in terms of making sure that the content that they're generating is relatively unique. It's not just on, it's not just about training on diverse data sets or training on open source data sets, say of like open source images where copyright may not be a problem for example.

Speaker

There are of course huge data sets that have been already created in this process even a couple decades back with data sets such as image net, which is to have like a variety of images across different tasks. But even using those data sets and performing different kinds of data augmentations, for example, to generate like other images and these data augmentations could be as simple as occluding or hiding some pixels on an image.

Speaker

As simple as just changing the color gradient of the images to like on green or grayscale and things like that. And it has been empirically observed that by making even minor data augmentations like hiding pixels, changing the color contrast and gradient, cropping pictures, rotating pictures, etc. These data augmentations have been empirically observed to still really improve a model's training process, not only in terms of extracting extracting more diversified patterns, and like more value-informed patterns on generating images, but also kind of just coming up with really new and like really new and other kinds of creative images on different kinds of augmentations that have been fixed to the model.

Speaker

That is probably the most clear answer that I can provide you in terms of how folks try to bypass these or. No, I appreciate that. I think that's really valuable context, especially around the, you know, going into the next here question that I have, which will be around biases. But before I do that, I am curious, a lot of these seem to be just inherently subjects that people, depending on who you ask the question, you might expect to hear different answers.

Speaker

So, at companies of scale, you know, big tech companies, how does the decision making work around that so just to give you probably a really, you know, thumb example that will probably never happen in real life, but say you're working on the LLM of whichever big tech company you can make your pick. And you found that it does or it generated a possibly problematic answer to one of your questions that you can see it doing again and again, like on a fairly common basis.

Speaker

Now, if you were to pull a sample of 100 people, maybe 60 would say that, yeah, that's problematic, 40 would be like, yeah, that's okay. Who decides whether that response is in fact problematic is it some shadowy leader that you never meet or is it you as an engineer, or like who makes who is the decision. Again, really important and complex question here in the mind. But yeah, it's quite important to be addressing this right.

Speaker

My personal experience in these large kinds of tech organizations which deal with training some of these foundation models. Work is kind of really divided into various different parts and handled by all kinds of different teams. When it comes to actually the performance of the models and what it's outputting, a lot of that work can go to the science or measurement kinds of teams that actually work on overall like model performances, say, even based on quality qualities, one of the most fundamental set of performance evaluation criterion, right, kind of generating say outputs and making sure if those outputs are good or bad or not right aside from using, like, aside from using open source data sets.

Speaker

Companies can also have close source data sets or their own kinds of data augmentations or generations that they might be using. There are, there is a whole suite of kind of teams and area related to rent teaming of machine learning models, which is related to actually having a simulated environment, where a model is spent with all kinds of adversarial problems to see how that model may well so that those kinds of things can in fact be corrected and kind of supplied back to the model for training process to make sure that these models are really resilient when they're pushed out.

Speaker

A lot of the work has also been open sourced by companies like Microsoft even I believe there's this open source framework called pirate. P Y R I T, which is this like open source, GitHub repository where a former colleague of mine has worked on, which a former colleague of mine is worked on for other users to actually create these kinds of different flows, these different flows to red team their model, or example. And so there are a variety of these teams that kind of either do like automated testing, scalable testing to get data set, we did back to the model for improving it.

Speaker

There are so potentially teams that can do more boutique testing of these models. Like a whole bunch of categories like even inside adversary, even inside the field of adversarial attacks, where a lot of the T, where some teams might even do boutique testing in terms of a human physically trying to say like break the model internally to actually learn different kinds of insights from them and trying to really creatively kind of construct prompts and data sets and other things in terms of really monitoring how good and bad the behavior is.

Speaker

There has been a lot of really interesting insight that we've gathered in terms of how the model deals with like data that may not be in English, for example, data that might be prompts that might be structured in the form of mathematical calculations, different types of prompts like role playing, for example, where you're kind of like steering the model to kind of like not do something good under the guard of like being a different person or something.

Speaker

So there's a lot of these like very interesting kinds of testing that. Now, with all this being said, I do want to acknowledge one really interesting point that you raised in a month related to hey, if one image is is not okay by 60% people but okay by 40% people right. So how do we make those determinations. Within that testing team right you're referring to that testing team being the person or the group that's doing the sampling.

Speaker

Yeah, yeah, there are two kinds of broad categories that I envision that this is kind of being dealt with. One, even by choosing typical models on like these large with these large like cloud computing platforms, there is kind of I'm sure a lot of safety guardrails baked into these models by default. To just be on the safe for site. That's a pretty basic but useful heuristic that it can compromise performance to some extent as well.

Speaker

But company it is possible that companies can make that determination at an initial level that that is probably still okay and not at the adverse cost of outputting something at that's one thing. Second, it's also quite likely where users have the flexibility to set different kinds of model parameters for themselves, not just related to how creative or deterministic the model output with parameters like temperature, for example.

Speaker

To not only pick different kinds of like evaluation metrics across measuring these arms, but to also define say like different kinds of like thresholds of tolerance that they might have two words specific kinds of outputs, because we really need to claim to be the experts here in terms of what's good, what's bad, so interesting culture. To some extent, we want to surface this level of functionality to the users in my personal opinion, where users can be empowered enough to choose those thresholds for themselves based on their culture and context and the domain of problem that they're doing.

Speaker

So far we do you think we are from that day though, because at least from my knowledge, I don't think this exists currently. There are a lot of like there are several not only open source tools, but also startups that in my knowledge have gotten created in the space that kind of offer tooling around kind of say, evaluating large language models across these different criteria and giving that power in the hands of the user. There are tools that do exist. Fortunately, a lot of this tooling however is also quite new.

Speaker

A lot of sorry sorry just to clarify I meant from the companies that make the LLM themselves like why is there a need for a third party to make this why can you know the Google's of the world make this feature themselves and ship it out. So, a couple of reasons why that could happen so one I think everyone is kind of making tooling around the space, not only the larger companies in the model vendors but also startups.

Speaker

There can be differences in terms of, of course, the the speed at which these tools can be built by large companies versus startups. That's one important factor in terms of getting more attention getting market cap to it can even be related to these companies are building tools that they might even say larger companies can tend to lock say users within the same platform. Or within the class of models, right, where they would work well, they would work great with models that belong to that same company but not necessarily other models and that can kind of vary, right.

Speaker

Startups or see open source tooling that gets created in the space that can kind of boast of a more versatile sort of perform a more versatile sort of usage across different models, so that users don't have to get locked into using a particular model from a particular vendor. And look at the entire span of models that exist in open source by these different model vendors and take a more informed decision for themselves in terms of how to actually bake in the safety practices in their models and how to actually mitigate the costs because this is also very expensive, right. So say users using smaller models, for example, or using open source models could have, say, even merits for state companies or users that are starting out very early in this race and trying to figure things out for themselves in terms of building applications,

Speaker

committing costs, and gradually scaling up their usage, right. So in terms of like what's stopping a company, a large company that's building these models to also build those tools, nothing. They're already getting created in the space. Startups are also working on it, which can be faster on an average and larger tech companies. And there are open source tools as well, which to some extent open source developers could even use for free for doing testing at some scale. Right.

Speaker

So there's different value propositions that are offered with free open source tooling for quick and dirty experimentation to say tools by standalone companies and platforms that ideally are platform agnostic or model agnostic to using large cloud computing platforms, for example, that boost hold variety of privacy and security guarantees enterprise guarantees, a lot of compute, right, a lot of the tooling for scalable experimentation and things like that.

Speaker

So these, there's these different like market sectors and different values offered by. Really cool. Yeah, I mean, this, this whole, like I'll use chat GPT, everyone's once in two days or whatever, but it's just mind blowing to me just the entire absolute universe or maybe universe says that exist behind, you know, just a generative AI as a technology and really what goes behind that so from all of what from my very again limited knowledge on this subject.

Speaker

I do think what you said about giving the power in the hands of the people that he let me decide how a woke or non woke I need my responses to be. I actually think that is, I don't at least I cannot think of a better solution to this because as you said, I do 100% agree with that that companies, I really should not have either the right or the power to make those decisions for other people. However, having said that, you can certainly go too far in the other direction. I'm referring to grok by Twitter or X AI, which to me it just, you know, feels like not a very good product because it just feels like a gimmick to me.

Speaker

It's just like a toy that people would use to generate stuff that's just offensive, because it's the only place you can do that. So, I guess my question to you, and again, I know this, this has nothing to do with, you know, what you do for work. I guess it does have a lot to do with that but where you work or you know who your colleagues are or whatever. But I'm just trying to understand from your perspective, where do you stand on, you know, just the scale of where you're trying to be inclusive and helpful, but also accurate.

Speaker

But what point does that water get muddy and what can we do maybe as not just people that build the technology but also people that use it to make sure that it's, I guess everybody can't win, but at least, you know, it's at least acceptable as a solution. That works for most people. That's a very broad question. So, it is. I do. I do agree. Yeah. So I can, I can try to answer this in kind of, I guess, from two perspectives, one from the user, and one from the AI practitioners who are actually developed.

Speaker

Perfect. From the user's perspective, one thing I would really like to emphasize on is user feedback. As you rightly pointed out, no one's an expert really at determining what is right and what is wrong. And so as these new models come, like get released and get tested by real life users on real life use cases, it's really the first waves of user feedback that really guided how these models actually have to be constructed. And I'm pretty glad that even though heads of large tech companies have sort of alluded to the importance of getting user feedback, claiming that their products are not perfect at all, but

Speaker

maybe backing on or depending on, say, initial user feedback to better understand how to serve them. So on the user side, of course, if the user sees something offensive and things like that, there are usually like functionality, of course, provide user things like that, or kind of, of course, say, like treat the parameters of a model as you discussed before on like what level of tolerance do they want them to have and things like that.

Speaker

Right. It's, it's a very interesting optimization problem where the user is trying to balance performance quality, safety, computation cost, latency or the speed at which results are served to the user. So there's these really interesting just like four buckets right off the back that I tried to get out in terms of how the user would actually do these, how the user would actually optimize across all these lines, right.

Speaker

I won't claim that there's a straight answer to this because this is really an optimization problem that I practitioners would have to figure out for themselves and for the users. So what I can see is that feel the evaluation tooling, or whatever tools and platforms that users are using, ideally needs to be built in a way which is simple and actionable for users to understand the results. So I'm going to draw some actionable conclusions out of them in terms of where do they want to see maximum improvements to their models, comparing different iterations of models across these different metrics.

Speaker

And what kind of makes the most sense for them, practically, to deploy as a solution. So in the sense, I feel like I would even argue and steer away from AI and say that this is actually can be dealt as a user experience question. In terms of that's the tools that are building brand new user experiences for a relatively brand new set of concepts, which find me foreign to most users, especially the users that are starting out.

Speaker

How do they make sure that user experiences are simple to latch onto are easy enough for people to, for people to enter and try to use them. How do we make sure that the surfaces where those tools are deployed as an entry point are kind of visible to users and understandable to try some quick and dirty evaluations initially, and then kind of imagine or visualize creating more complex workflows and evaluating across these lines. And how do we make sure that the results are actually very clear to the users in terms of where their models are good, bad, how do they compare against different models, against different parameters, hyper parameters, and what kinds of actionable insights can be explicitly provided to users, them to go back and improve their application based on those results.

Speaker

That's the answer to like how a factor to do this, except like perhaps say the tool makers making sure that the tools are easy to understand, and for people to start using them and start onboarding to them more and more, so that these decisions can be taken iteratively and it is developing as hard as this obvious be an iterative process. There is for sure around that. It's almost what you just laid out almost brings to mind about how, at least in the early days of cyber security and this is still probably true to at least on a fundamental level, but it's the expertise of the people doing the hacking that actually leads to more robust systems. So if you apply that same analogy here it almost sounds like it's the as the adversarial prompts get better.

Speaker

In fact, in various, you know, it will trickle down and make the models also get better and it'll almost be like a, you know, adversarial relationship but one in one in which at least this particular, as you said AI practitioners, they would benefit from that interaction. It's just the balancing act of that, you know, of the power dynamic that which I think will be really helpful for, you know, for me to witness and for you to build.

Speaker

Well, you should have thought of that before you, you know, started doing this or living on this absolute cutting edge of technology that you're in currently but, you know, generally, so just so blown away by all of the not just the expertise that you bring but the way you have this, you know, breaking down this really complex questions, like in my head, a lot of these questions don't even make a lot of sense, and I'm just throwing them at you and yet you're able to somehow cobble together a very well informed very articulate response so really appreciate that.

Speaker

And the really last question here for you at Vithya before I let you go. How far do you think we are from AGI and for our any uninitiated listener, if you could also break down what AGI at least means for you, that would also be really helpful. I'm going to be very honest with you. So first I'm going to talk about, I guess, what I think of AGI and then how far, yeah, maybe from in things like that. So if you, I, in my mind, I personally decided to not have an opinion about what AGI is.

Speaker

The reason for doing that is that I, there's just so much information overload that is happening in today's word with how fast this industry is evolving, that before anything I just want to soak in all this knowledge and kind of try to capture trends and where things are rather than jumping into opinions or conclusions. That's just been like a personal sort of like priority of mine in terms of how I want to look at this. At the same time, I have seen that, of course, AGI has been defined, say by companies like open AI and others as like a really intelligent system that can

Speaker

perform humans or overtake humans in most economically viable tasks. I feel that definition makes practical sense to me. I'm not sure if that's a comprehensive enough definition because just doing economic tasks isn't the, isn't like all that human intelligence is about there's a lot of challenge to that. But it makes practical sense to at least start out with, because that is how these model lenders are likely going to get their money back or actually earn profits from the usage of these models, right?

Speaker

A lot of these models because they're sold to enterprises. It only makes sense that if these models are useful at economically viable tasks. That is how everyone is going to earn a profit. So, I feel it's a good definition. Where AGI is going, how long will it take for us to get their very opposing perspectives among like different. So, how important is the idea for them versus not? How far is it going to be? How do we even have the problems like climate change in the scope of like as we kind of develop this AGI and or kind of like develop models that have like really high electricity consumption and things like that, right?

Speaker

So, I kind of also wrote about this in one editorial that I had written in like a separate kind of news reporter a couple weeks or like a month back, where I said that I kind of am not even inclined to distinguish right from wrong in terms of who's right here or how much time is going to take. But what I wanted to point out is that AGI is a fluid concept number. This is really the overarching idea of my answer that I'm going to expand upon right now, that AGI is a fluid concept.

Speaker

And because of that, there is pretty much no defined sort of timeline on like when AGI would arrive, or how do we even define what level of AGI has arrived at what point? The reason for that. The definition of AGI, I believe has also changed over the years, even by these model vendor companies and things like that. If you consider the definitions back in 2018 or before to now and things like that, these definitions have changed.

Speaker

Also, as new models came in and as these large language models came in, understanding of human cognition has really improved in terms of distinguishing between what intelligence can be generated by a computer versus what is truly human intelligence. Even then calculators came in like say before calculators doing mathematical calculations was a very important tenant of say measuring human intelligence. After calculators came in, we could think about a broader set of problems to deal with.

Speaker

And so the definition of our understanding of what it means to be human, what it means to be intelligent, and the different branches of cognition and understanding about the human brain has evolved and improved over the years. And so as we look at these large language models, how versatile they are in terms of generating content in terms of writing of language with perfect grammar, for example, in terms of generating even like poems or like things that people deem as creative, but even humans at the back of their mind would take inspiration initially from somewhere and then come up with something and things like that.

Speaker

So our definition of intelligence cognition, and what it means for a system to be general in all senses of intelligence is fluid is evolving. And so because of that I feel as these new models come in boasting new capabilities, like our understanding gets refined on like what can these models do more importantly what can these models not do. These models like GPD4 or whatever they can answer really intelligent questions really well grounded in real world data all that these kinds of models are likely not going to be able to answer if 3307 is a prime number or not.

Speaker

Oh, interesting turns out really yeah turns out that's not the kind of past that it's been built upon. I have seen examples where that's so interesting I had no idea I just assumed it would be okay doing that. I have seen examples where these kinds of large models and I'm definitely not looking to back bounce these models by any standard. Of course, yeah, learn about what these models can do cannot do. And this is all really helpful feedback even for the model vendors and everyone moving forward as we all figured this out by ourselves where I saw this one prompt is 3307 the prime

Speaker

response was 11 can divide it. And this is the answer with the decimal points and hence it is not a prime number when in fact I believe it was kind of big right. So a lot of people didn't talk about it. One of the conclusions that was raised was that this model is a next word predictor. It is designed necessarily for doing mathematical operations of this type or the kinds of different neural and symbolic elements that could be required to truly understand and operationalize mathematical calculations of this kind is probably not encoded in the model just yet. And so again coming back to your answer.

Speaker

What I do cannot do. How far away are we from different kinds of intelligences, our understanding of it. And so the overall definition of AGI towards the economic tasks initially and I wonder what even comes next after economic tasks. There's got to be something more than that I feel that this is this is going to be an evolving field. Obviously you're going to see a whole lot of developments in the next five to 10 years.

Speaker

I believe there will be lots of lots of developments around like multi LM setups and agents, and these more complex workflows to kind of automate bigger parts of like enterprises and bigger parts of the processes. And we'll gradually get to expand the scope of where humans actually offer value and how humans kind of use these tools. So cool. That's just, that just blows my eye. I actually used to be convinced that now, like the times that we live in are in fact the best time to be alive. And then this whole thing came about with generative eyes and now I'm like, if there was even a little shadow for doubt in my mind, it's gone like I don't generally I don't think there has ever been a more exciting time to be alive.

Speaker

Not just in terms of technology, but really what it enables for everyone around the world in so many ways. Technology is in fact such a great level and the ability to leverage it to do what you want is just an ability that just simply didn't exist even 100 years ago from now. As I mull on that and, you know, the, I don't know 31 other things that you've left with us here today. I'd like to thank you so much for taking the time.

Speaker

Or, you know, not just going over and sharing your expertise, but you're deeply rooted in sites in terms of not just the work you do but also what your own feelings are towards all of these really con, you know, complex topics. And yeah, I think it takes a lot of, you know, a genuine interest to have this amount of brevity on these complex topics as such as these. So really, really grateful for your time and I cannot wait to, you know, remain connected and see all the amazing work that I'm sure you'll do in this.

Speaker

Having, you know, for all the kind and encouraging words. Yeah, I had a lot of fun in this conversation and I really love the thoughtfulness with which you kind of both really complicated and understanding challenges and questions in terms of how we want to kind of look at these things. And that kind of, of course, like also motivated me to think more deeply about these things literally right now in this conversation.

Speaker

So really appreciate kind of the thoughtfulness that you kind of brought to our conversation. I really hope that some of these converse, some of the stuff that I said, some of our conversations, hopefully provides a similar, a similar amount of value to your audience as the rest of your viewers and podcasts, podcasters did as well. And yeah, definitely like looking forward to kind of also seeing like your next set of podcasts and the next set of people you get to talk to.

Speaker

And what I can also learn from them. But yeah, I hope you have fun in this conversation and I hope that the audience watching this also also has fun. That brings us to the end of episode 29 of the ready set to podcast. Thank you all for sharing these conversations with those that continue to benefit from them. If you would like to support me, the easiest way to do that is by subscribing to the YouTube channel and or leaving me a pro five star rating on Spotify. Catch you all in the next one near episodes every Wednesday.

Transcript-backed moments

A few lines worth stealing before you hand over the full hour.

00:00:00

You can think about responsible AI as a big umbrella of safety guardrails in that you want to bake into a model across different pillars like bias, transparency, privacy, robustness, interpretability, etc.

00:00:12

If you were to pull a sample of 100 people, maybe 60 would say that, yeah, that's problematic, 40 would be like, yeah, that's okay. Who decides whether that response is in fact problematic? Is it some shadowy leader that you never meet or is it you as an engineer or like, who is the decision maker here?

00:00:29

I got to work with machine learning systems research to AI for accelerating IT operations to even responsible AI in my final rotation. Maybe I should have led with this, but where does responsible AI come into what you just described?

00:00:44

How do we make sure that whatever the model outputs actually grounded in real-world data and is actually correct? How far do you think we are from Agi? Welcome to the Ready Set Do podcast, where we learn from journeys of not experts who are just two steps ahead of us.

00:01:06

I'm Naman Pandey and in this episode featured not expert is Advaitya Gamavan. Advaitya has the incredibly unique role of working on responsible AI at Microsoft as a machine learning engineer to more so than the other episodes that I've put out here.

Show notes

Advitya works on Responsible AI at Microsoft as a Machine Learning Engineer 2. I am extremely stoked to present this incredibly nuanced discussion on the ethics and responsibilities that come with building mass-market generative AI tools. We discuss the technical foundations of Advitya’s career that brought him to this cutting edge of technology, and also go over how GenAI tools by big tech navigate critical concerns such as privacy, ethics, inherent biases that creep in all the time; and the aspect I was most curious about - WHO is the actual decision maker when it comes to flagging something as problematic and/or offensive.

More in Build It

Same mess. Different guest. Pick the next conversation that feels closest to your real life.