Episode 124 transcript
How To Break Into Site Reliability Engineering on A F-1 Visa - w/ Sai Joshitha transcript
and a Kubernetes pod just quietly died. A guy in a different time zone taps his card for coffee and it works.
Questions this transcript answers
Search-demand and topic-shaped questions mapped to the exact moments where this conversation answers them.
Full transcript
Timestamped transcript for How To Break Into Site Reliability Engineering on A F-1 Visa - w/ Sai Joshitha, with answer markers attached inside the conversation.
I kind of applied almost 200 plus applications every week but I didn't give up. That's what I can say. I didn't give up. I I tried every day. Every single day I tried. That's Sai Jooshita, senior site reliability engineer at Visa, the same company that runs the systems behind most of your payments. By the end of this episode, I promise you will know how to land a role that most engineers don't even know exists. That's not how interviews goes for SRS that most of the interviews for SRS is how you will handle if the system goes wrong. We do have one technical rounds coding round but it's not like lead code things. H so no lead code. Well, so then what are they actually measuring? It's not like can this person write the code? No. How this person can handle this situation? She lays out the exact tool stack recruiters screen for
and the one thing she wishes she had started back in her masters. If anyone like kind of scared of coding, I can say very helpful if you wanted to become an SR tool, right? It's a very easy to learn at all. Join me in welcoming Sai Joosha to the ready set to podcast. Without any further ado, let's get into it. Josha, welcome. Uh, hey. Say hi. So excited to pick your brain around a particular topic slash field that firstly I actually was not even aware that this is you know an entire direction for taking one's career. So for those of us that are not super familiar that includes me what does a day in the life of a senior site reliability engineer at a company like Visa look like? Uh so yeah it's a kind of um so it's I I'll just uh talk uh generally what exactly is you know honestly honestly and actually that's a very good
question. So when I tell people uh I'm an S sur uh a lot of people you know ask like what even that means actually many people know about don't know about this actually. actually. Yeah. So the simplest way I explain is that uh you know my job it's kind of help make sure uh critical systems uh remain reliable um available and predictable. Uh so but the day-to-day work is actually um you know very uh different from what people imagines um kind of you know some days I'm looking at infrastructure health some days I I deploy things I monitor and uh I automate and I one thing I really uh like about uh reliability uh engineering is that uh no two days are exactly the same. Uh so we always we always like uh need to be available uh for everyone like 24 by7. So like you know if you are using some applications you know the
outcome is like whenever you are using something you know only uh you know okay the for example if you started a transaction you know your transaction went in and you know you are able to see your balance and everything but you never uh think about like how this transaction went and how this is this was successful right successful transaction uh so every S sur will be the part of it like uh we kind of uh support every application behind the you know every successful transaction something like that we support and we make sure our applications are uh available for every customer and every um you know people out there whoever are doing this transactions that's what uh my you know my day-to-day um kind of roles and responsibilities. Yeah. Yeah. Yeah. that that's super helpful and really appreciate you sharing all of
that rich context for us to you know just ground the rest of our conversation in. You mentioned that sometimes you'll have to attend to incidents. So what would be an example of you know an incident? It doesn't have to be something real obviously but just for us to contextualize um where what is the site what is the incident in question and maybe where do you come in in terms of taking action to as you had said making that system more resilient. resilient. Uh so we kind of get many many incidents because uh we kind of operate uh you know kind of very big applications. So kind of uh I can say some example like some simple payment example because uh that's something most people can relate to. Um like imagine you are trying to buy something online. So you you add items your card right so you enter your kind
of payments and information and then click pay that that's what uh we do even I I do the same. So from the you know uh customer's perspective uh the the incident is very simple like I tried to pay and it didn't work. So but behind the you know scenes uh there are many uh systems involved for example our infrastructure there are many systems uh so we we kind of like have lots of systems and lots of um you know APIs so there may be an API or database or some you know authentication services payment uh processings uh services kind of network communications and um uh several dependencies all working together right So if any of those components start behaving unexpectedly of course customers can start seeing uh issues right. So as I mentioned uh we have many systems and applications. So for example if there is an any issue uh kind of in
networking side we definitely get an incident. So and also like we we don't immediately know that right that issue happened and some problem where the problem exactly happened. So we just know like something isn't uh like uh behaving uh normally. So the like maybe we are seeing more failure transactions uh than expected or maybe response time suddenly increased or maybe you know s some servers crashed or maybe some services time dot the first step we usually understand like what's actually happening. So that this is where reliability engineer typically comes in. So uh we are looking at monitoring data logs alerts uh recent changes and trying to build a picture of the situation. So what changed what involved in the last few you know days or last 24 hours that's what like uh we do. So if an incident comes that's what
we take take up that incident and we check where that went wrong what exactly happened in the last 24 hours. We check logs, we check metrics, we uh kind of try to troubleshoot based on the incident what exactly happened, why that transaction never went and why you know that service uh response time taking more time. So that's uh that's what I wanted to say like it's kind of um I can say that one uh ser service is struggling because traffic is much higher than the expected sometimes.
Yeah. uh in shortterm kind of some code change changed something due to some deployment. So that's what incident occurs and some that's what how issues happen from our end. So it's a kind of it's very you know we we need to be available uh like every time to taking care of that incidents. If we don't uh troubleshoot or resolve that in incident, we definitely see some high failures and customers might uh you know uh see some uh loss from their end. Even for for example like if we consider like if you are doing some transactions if your transactions went through you definitely check your bank or something like that right. So that's what we are here to avoid that such scenarios and even if we have that scenarios our responsibility is to um not to repeat that incidents and something like that.
That's awesome. Yeah. I mean I actually it makes sense everything that you're saying I'm I'm now like yeah there this is obviously somebody has to do this and I I can also see how when you had said earlier that no two days are the same. It's like a new challenge every single time. I see where that comes from because yeah, like there's so many moving parts here that yeah, any of those if they if they fail, step one is to obviously figure that out, which I can't imagine that being easy already, right? Cuz there's just so many systems involved. involved. But then from there, you have to do like the root cause analysis. Why did it happen? How to make it not happen? It all sounds super super fascinating. So I am curious from there if we were to shift gears a little bit take us back to around the time when you were trying
starting to apply for your masters. Is this something that you knew you were trying to get into? Like was this specific you know this type of uh you know um investigative tech work if we can call it that always your calling or was it more of just you know something that just kind of fell into your lap? Oh yeah interesting. So actually uh I started my masters like everyone like every engineer I actually didn't plan it uh plan it actually uh honestly I I don't think I had this entire path planned out so so uh kind of if you had asked me during uh engineering whether I I would uh you know know end up working in uh reliability engineering uh cloud in infrastructure and uh distributed systems. I probably wouldn't have predicted it. Uh like every engineer I also focused on coding, writing, programming and everything. So it's it's
always a curious like every stage of my career uh exposure to a new set of questions. I started with a software engineer and I build applications and then I became like curious about you know kind of what happens after applications are deployed. Then I became interested in infrastructure and operations uh then coding. So then it started like cloud computing then relability. So for me it wasn't one big decision. uh it's kind of a series of uh smaller decisions and out of curiosity then it continued like uh reliability like kind of um I mostly wanted to uh do it and that's how ended up as um focusing on more reliability side and S sur side and that really uh you know I I genuinely loved um how I'm working today and um how um you know I'm able to have exposure to different technologies is as an S sur. Totally. So for our listeners, can you
uh share where you went for your masters and what you studied and maybe even cover any internships that you did? Just really kind of looking for a brief overview around your master's experience itself. itself. Oh, sure. Yeah. Uh so kind of my masters I studied uh in University of Michigan Dearborn. uh I have um completed from I mean completed in the year of 2023 um in masters I have opted for computer and information science so yeah so it it's kind of interesting uh because I I actually this was also not planned at all it suddenly happened uh but I'm I'm very lucky to uh did my masters because you know before masters it's only kind of you know uh I my exposure is kind of very you know local uh like within my area but after coming to the US I really got a good opportunity uh for more networking side and more u exposure to technologies and
how you know the education system different from how I completed during my bachelor's so yeah the uh so my for your question I completed in the university of um Michigan diaborn in the field of computer information science. So I haven't uh during my master I haven't done any internship but after my uh masters I had an opportunity to work as a cloud um intern uh in the one of the uh in the company uh as a cloud engineer intern. So that's where my uh career started in the USA. Got it. So I'm curious from there because of what you just said, right?
Like it's it sounded like you said you were not necessarily holding out specifically for you know reliability engineer type roles. So when was the like I'm just curious about like your job application process itself slashin and all of that cuz maybe this is just because I'm on look you know from the outside looking in but yeah I I struggle to kind of connect the dots between having generalist computer science experience which is what you said you were studying slash you know upskilling for and then this almost super niche role that you're working here today. So, how did that bridge appear and kind of can you walk us through that? Maybe, you know, the application process, the interview process, just anything for somebody that's looking to walk your footsteps that's just a couple steps behind, you know, that they can follow
and watch out for any pitfalls and such that you came across. Oh yeah, sure. So actually I had an um experience uh before I come uh come to the masters as a devops engineer at um at os company. I have completed there um you know 3 years as a devops engineer. Then I I planned my mastersh here in the US at University of Michigan as a computer information science student. So honestly it's uh not at all easy uh to get an interview uh due to the job market as everyone knows I think still we are struggling with that. So even I struggle a lot but um so if I wanted to give a suggestions to whoever is trying you know out there like I kind of applied almost I 200 plus applications every week that's how I yeah it's if we what I observed is like when I apply more I usually get interviews uh in that way so that that's what I observed that
pattern then I started applying more and more applications then even like everyone I I I got um interviews few of them I you know I reached all the last till the end um so that that's how my interview uh experience was like I almost tried few for few months then I got a opportunity to work at um as a cloud engineer after trying multiple times but I didn't give up. That's what I can say. I didn't give up. I I tried every day. Every single day I tried. So yeah and here networking is also very important. So I one of the thing I did is emailing the recruiters.
Every applications we have every recruiter I I emailed I contacted them. I provided my uh resume. That's how I ended up landing in uh cloud engineer. Then up then I I still wanted to uh you know work as a um deops and s sur even though I'm working as cloud engineer because of my interest. Then I started again parallelly applying for full-time. Um that's how I ended up in the visa. Uh it's a this is also like it's a kind of I applied and then my resume got picked then I got I went into interview then that's how I ended up in the my company of fintex. Yeah. Got it. So when you were so if we were to just focus on the strategy piece because if you're a computer science mast's degree holder there's I feel like at least I can already think of 10 just completely different job titles you know that you could be applying to. So
obviously you have your SDE type roles you have your more specific machine learning engineer type roles. I know a lot of computer science graduates that apply to data science type roles. Some of them apply to business analyst type rules. So, right. So, how did you think about your application strategy itself? So, I know you said 200 per week, but did you have a strategy or plan around what you were wanting to go for or was it just a pure spray and prey approach?
Oh, yeah. Uh I I know I clearly know that I'm going to apply for site reliability engineer or cloud engineer or software engineering. um because um I I feel like these roles have more u you know uh exposure to different technologies for example right now I I work on different different even though I'm a s I have got I got an opportunity to you know work for to know about the different technology that I previously don't know actually so so I I I know that I it's kind of um I did not kind of apply to each and every role I I know that I I'm going to uh have only S. So yeah, my resume and uh everything is set up for only for SR roles and cloud engineer and software engineering roles. Um yeah, it's it's it's not like I I applied to each and everyone. So I I have planned it. I have planned everything and uh as I mentioned I I
know if if my resume is uh something suits to one job description I kind of did everything what I can do you know reaching out to the recruiter emailing to them and providing my resume and following up everything I have done each and everything that's and for each and every application it's not like our resume will source to every applic I mean every Sorry roles description they have mentioned but still uh every uh you know company needs who knows about how to check the monitoring how to check the logs even though different teams you know different company uses different tools for monitoring or you know for cloud one company uses GCP one company uses AWS it's not like we need to have each and every every technology out there in the uh resumeum uh I we need to know like uh even though if we if we we need to be ready to opt whatever the
technologies out there even though if I know AWS I I should be able to know the basic things that I can also do that in GCP side so that that's what I I learned the basic things and I even though my resume doesn't suits for some role I applied I tried I tried to apply it's not like they are only looking for you know particularly these things they they're looking for the things that how how we manage to work or even though uh they are using different tools that that that's what I think and I whatever we have the application that uh naming s sur I just applied to each and everything uh I maintained one resume for one s and cloud engineer and devops I included whatever I know uh the toolings and all uh and I applied that that's what my strategy nothing else so instead of applying to whatever role just um decide what exactly you know and
what exactly you can perform then target on those roles that that's what I did then it it worked yeah no that makes a lot of sense when it comes to the interview process itself for S sur roles what usually does that entail how many rounds are there what are some of the questions that you know are must nos and then maybe because you prepared for these if you could also share any resources that you remember maybe any like books or YouTube videos or just yeah really any resources that really helped you and that will help anybody preparing for these roles uh that would also be really helpful. Oh yeah, sure. Sure. Actually like yeah even when I was applying uh I kind of looked for some resources like who could help me out because it's a very new to me here here especially in the US.
So yeah uh definitely why not? So but uh what I did is see we have multiple technologies. It's not like we should know everything as I mentioned the for uh before uh one example we have different clouds right AWS GCP and Azure so learn at least one uh you know everything like completely each and uh end to end so it's it would be helpful for for everyone like even though if if they got a project on GCP side or Azure side it would be helpful for them so tools and and The resources I have used for applying I mostly used LinkedIn for sure. I think everyone uses that for still till now. I think uh I have uh also uh checked in indeed um um I think I I didn't remember that uh but I have checked one more thing which is um most of the startup companies uses that I really don't uh glass store maybe yeah sorry I meant
more so in terms of upskilling actually like when you were trying to prepare for the interview itself I meant resources around that piece h mostly uh you know it's a kind of For me one of the uh advantages I have from uh previous experience that that worked and uh one more thing for upskilling I have used YouTube channels whatever it's not like kind of particular channel I followed followed is understandable for my case like for me I have followed that YouTube almost I have used Udemy also for most of the courses which I don't know which I like and um masters helped me uh I have you know a few courses where that focuses on cloud and um software engineering subjects or something like that that helped me and the projects helped me uh uh even though see uh every everyone uh can learn from every video what I did is
even if I'm learning right I plan to do hands-on uh in my local system that worked actually uh mostly Um but yeah YouTube and Udemy the courses that that helped me most and uh I uh and also the as I said one one of the advantages I have is like having a previous experience um that also worked uh for me that that's what I have uh ended up and interview rounds and all right yeah right as I think everyone knows like we have different interview rounds um yeah technical technical the technical so yeah S is a kind of interview course like you know um mostly conversational based uh it's not a technical so okay tell me what is um uh you know what is that that's not how interviews goes for SRS that most of the interviews for SRS this how you will handle if the system goes wrong that that's how they so there's no lead code rounds and stuff
like that uh uh no Yeah, we we do have one technical rounds, coding rounds that but it's not like lead code things uh mostly for automation side the basic programming we we should know at least uh and as I said uh it mostly we we have to know like they'll think they will check whether we are able to handle the systems if they went down if they go down like something like that they check yeah yeah it's not like can this person write the code no actually that that's not they will uh look for the candidates the they will look for how this person can handle this situation because that's how S right so the interview always based on the conversational bas based like how you would u you know troubleshoot this one how you would think if if if you are in this situation that's how conversations will go uh during interviews for an S sur
sur got it and then Is there like a technical design round or I guess and maybe as you answer that something that would help also is in terms of just skills or tools specifically or maybe languages that are maybe must haves cuz I know when they're asking you these behavioral questions. Yeah. you're you're obviously you know replying from your past experiences and such but I'm sure some somewhere right in the back of their minds the interviewers are looking for like does this candidate know a certain skill even when it comes to like cloud infrastructure right is it more docker kubernetes I guess what is kind of the basic like musth have skill stack when it comes to s sur type roles so for an s uh right now everyone knows AI is kind of booming. So I don't know the recent interviews but when I have the interview I'll yes uh I definitely
you know uh they asked me about how my cubernetes knowledge is like because every company will host their application on the in cubernetes right so they definitely check whether we have the kubernetes knowledge cloud knowledge do we have and also as an s one of The most important thing is like monitoring uh we call it as observability. So so if you if I say the observability uh we need to know how logs works, how metrics works, how something we need to know how create the dashboards and all. So because see every we have every systems uh running or multiple systems. So we have to know where the things are going going wrong. So for that visibility we need the visibility right. So I think for every SR interview we definitely uh they definitely ask whether we have that uh observability knowledge which is like
dashboards like we need to know tools like graphana kibbana we have such tools for logging and metrics. So every S sur should know that actually at least these are the primary tools um they'll definitely check for um we have different uh tool links but for me as for like uh I they check it for me like whether I have graphana and primus kibbana elastic knowledge or not uh for the in the cloud side do I have at least one cloud knowledge and containerization as you mentioned yes of Of course we have to know the cubernetes they'll definitely ask that and um and also like for automation um see for example for as an S sur I can say we have so many manual things not only whoever is doing the S role definitely have so many manual things. So automation is uh needed for in when it comes to the coding part at least you have to know uh scripting uh
or some uh Python programming uh to survive in the automation side because it's it helps because we cannot do repetitive task continuously right so yeah that that also works in programming side at least you should know one programming knowledge uh you should know containerization side kubernetes docker and the cloud cloud side you should at least know at one cloud and monitoring side you should know at least three at side you should know at least three tools maybe uh like as I mentioned graphana and kibbana primeus and all so networking tools you have to know how networking works how TCP and all that also we have to know uh yeah the these are the basic things that every SR should know about it uh so and most important even though We know that we we know the tools right even though they check if they if we know the tools
they definitely check whether we have that troubleshooting skills. So yeah not only the tools uh we have to know the how we can troubleshoot with those uh using those tools because the tools itself won't um you know solve everything. So yeah, we have to know how to use those tools that that's what they tested tested during interview. That's Yeah, that's awesome. I love that. I feel like that is so comprehensive in terms of anybody that's trying to prepare will now know exactly where they must have or where they must be at least in terms of the baseline of you know like if you don't have these skills like let's change that right like let's just quickly teach teach ourselves at the very minimum these things and once they do get all of these up to speed I'm sure there will be other technologies that they just come across that will make
sense based on the role based on the industry based which becomes easier once your fundamentals are pretty yeah I agree about that but but yes uh you know problem solving is actually worked for me throughout my career that that's what I can say knowing uh tools and technologies is different and knowing how how we uh knowing how to use them is totally different uh so Yeah, learning is different uh like uh how to uh you know learning how to use those to tools also very important things actually that actually worked for me instead of knowing tools uh what exactly worked for me is knowing how to uh you know solve the real time problems and all yeah that that's that's what I it worked for me yeah I think for yeah for every uh role that the most important thing that For sure. Yeah. Yeah. I think it's such a great call out
because Yeah. It's one of those things that you know this, right? Like you've been told this before, but it's easy to forget that ultimately you're there to solve a business problem. Like you know, you're there to either increase revenue or lower costs. Ultimately, every role across every industry in any, you know, uh job title boils down to doing at least one of those two things. And it's yeah, I like that you said that because yeah, it's very easy to just forget that and you're like, "Oh, I'm this is this and I'm do I'm doing that that that but no, you're actually just you're either doing increase of revenue or you're doing decrease of cost. Everything else is just Yeah. you know, an abstraction to to that you put on top of either of those two things. So yeah, that that's pretty cool. Um you mentioned on this or you
brought this up briefly a little bit before but around the AI services that Visa is working on I think they're called Visa AI as service what I guess what can you share around that piece how has your job changed as or because of that new you know direction that visa has taken and what does that mean for the future of the S sur role in general? Uh so yeah here when I say um AI uh implementing AI in our uh task so I can say you know uh we actually manage different uh services and different uh infrastructure on prem and on cloud. So uh it's not like we don't get any issue uh so because see we are uh using some very large infrastructure distributed systems. So it is intend to uh have problems like one day and another day everything is silent. So so here uh and one more one more important thing is as an S sur every uh team will have on
call members. So yeah so for example if you are receiving one alert for example just say some some Kubernet is crashing. So here this is a kind of repetitive uh issue. So every time a on call member goes and checks why this is it's a repetitive one right. Yeah. Yeah. So here uh involving AI is a kind of removing level one of troubleshooting things like every engineer no need to go to the cluster and uh you know resolve that why pod crashes and need to do manually restart that pod and all. So here we if we uh indulge AI here AI can also do that right it can um restart that part and it can uh you know scale up and scale down everything it can do.
So here uh that's what it means. So due to uh I mean we as as we have many manual things here uh things like uh the most priority uh uh things are like availability like scalability and monitoring performances and operation systems like uh you know those all things can be also maintained by AI but not everywhere in at small uh at at one extent it it is good like um as They I mentioned one one of the example that kubernetative board things. In a simple way um AI can help in few of the things like checking logs what happened uh at the back end one job is failing. It can check the logs and it can tell the town call members this is what happened and and it we can also train uh train it like um to troubleshoot that and uh uh resolve that by restarting that that basic things uh AI can perform but uh not a big problems because see it's a
repetitive one that AI just need know the context whatever we provide it's it cannot um you know uh um troubleshoot every other problem uh every day because production is kind surprises every time and every day uh with new issues right so what I think is like AI can help at one extent kind of I can say AI and humans can be partners because yeah that's awesome really appreciate you sharing how AI has already impacted and will continue to impact even in the future about and just generally around that s sur role as it pertains to that on on a similar note uh if you had to share one piece of advice for somebody that's right now trying to apply for/in for S type roles what and you you're only allowed to say one thing right so what would be the one piece of advice that you would like to give to these future s uh professionals uh
so if I really wanted to give advice to the SRES so I would definitely they do networking um and learn as see it's a s sur is a combination of different technologies actually it's not like see for developers it's a a kind of focusing on one different programming long ways like and if you know long ways and it's it's easy to write the code and end to end and everything so as an SR I feel like we we should know different different different things and uh and all like as I mentioned And uh we use cloud, we use monitoring tools, we use networking tools, we use like kind of thing. These are the things like C we we involve in CI/CD everything each and everywhere we we touches actually. So if you want to be an S sur I could say you can learn you know every not not every tool but but at least uh each tool from
monitoring side each tool from cloud side each tool from you know CI/CD side each tool from networking side that really helps not only it helps for SR it helps how system is working uh behind uh uh you know after deploying the code If see uh it's a kind of helps not only for SRS and for devops also like it helps like it's a multiple tools combining together we deliver actually that's what DevOps and SR are so I can suggest like learning different tools is very helpful instead of focusing and I could also say if anyone like kind of uh you know like kind of scared of coding I can say like it's very helpful uh if you wanted to become an SR just tools right so it's a very easy to learn and all so I I can um suggest like just learn uh different tools out there um we have different resources actually like it's very easy to learn also like that
for sure yeah I agree for sure yeah so that that that's what I say like I that's what I suggest do networking that actually uh see during my uh masters or uh before the joining here uh as an S sur um my I I still think like maybe I should have focused more on networking um more on u you know focusing building my profile that that that I should have started uh since my masters that that's I I really wanted to give this suggestion start building your profile uh that really works wherever you are and whatever the situation change you are in that that really works that that's I realized it late but yeah that that's really helps whatever you are doing um and uh yeah if if you really wanted as if you if anyone wanted to have uh wanted to become an S sur yeah yeah I really suggest to know about AI and the tools out there uh basic
programming that helps landing on as an SR Definitely. Yeah. It's one of those, you know, it's just a quote that I like to say anytime somebody brings up networking is that the, you know, really hard thing about networking is that the best time to do it is when you know you don't need to do it at all. And you know, it's it's just kind of the rules of the game which makes it so hard, right? Cuz yeah, I'm good. I have a job.
Why would I network? Except suddenly you don't have a job and then you're like, "Ah, crap. Why didn't why haven't I been networking? So, it's just cruel like that sometimes. Um yeah. Well, this has been so cool. Joita, thank you so much for sharing your experiences and kind of the amazing work you're doing. like truly thank you for keeping our transactions safe and functional and making sure that whenever I'm trying to buy something off Amazon late at night on a Friday that I definitely do not need that transaction goes through successfully without any issues. So, uh, yeah, me and all of my listeners, uh, extend a big debt of gratitude to your profession, obviously your team and all teams such as yours that do this really incredible, interesting, and as you said, never repetitive work cuz every challenge is
new, every day is a new challenge, so to speak. So, yeah, truly, thank you so much for taking the time here today. I'm sure anybody listening that's trying to get into one of these roles is walking away having heard this with a much better plan of action when it comes to converting and cracking these roles. So, thank you so so much. Yeah, thank you so much for this opportunity. It's I really enjoyed it. Thank you.
