Subjects computer architecture

Cpu Performance Intro 63F0D8

Step-by-step solutions with LaTeX - clean, fast, and student-friendly.

Use the AI math solver

Question: Return the User: https://www.youtube.com/watch?v=BflEUNBD1kU&list=PL-Mfq5QS-s8iUJpNzCOtQKRfpswCrPbiW&index=2 orignal youtube link all right so hope you hope you ever started a good semester so far enjoy your weekend the first weekend before this door and then so I managed to upload the primary resources on your wiki page so just go to the course website and it's going to directly to this wiki page so our section is e 221 P and so by within 20 hours of each of the lectures I'll try to post new lectures here they're mostly from the publishers you know slides with minor tweaks perhaps what if something changes why I'm presenting I might update them later as well but mostly they're similar they're trying to confirm with the reference book as well so this is the high level agenda of the labs your first lab would be next week Monday so check your registrations you have you Mike registers in either of the labs I know lab a 0 1 or 0 0 2 or B we have limited seats there so make sure you attend your own specific labs because when you submit your grades you need to submit to that a specific pool spend some time to read the rules at home I'm just going to showcase you the you know the website so gradually by each which I'll try to make these assignments available so the first four as I mentioned before are before your midterm there is one make up before the midterm for them for the 1 to 4 so if you miss with an emergency or anything else that might happen sometimes for us for the first 4 you have an option to have a makeup lab and then if you missed one from 5th to 8th after the midterm you have another option for another makeup so the first 4 labs we will use that RVs simulator so these are these are the links and the manual make sure you spend some time this week before your first prelab which is next week Monday to learn about the overall view and the manual and an after midterm we're gonna use the very like simulator which is this he comes so it's a compiler for very long so the first four are the assembly using the simulator and then the second for Rd the very luck one we we were just assigning your TAS today so by tomorrow or maximum on Wednesday I'll post an another notification regarding your TAS so each of the labs you have two TAS available so for lab a you have two Thierry's for lab B you have 2 D 2 another 2 TAS at the lab session so the first few weeks I'll be physically there as well in case something was going wrong or we needed to change some permission or you know because we are dealing with 70-something computers and on a submission system you need a chair there's also two over there so about the labs so you have 90 minutes free lab so a few days before your first lab or each of those labs I'm gonna post your pre labs so it's a set of instructions it's a PDF with a set of instructions that walk you through what you are expected to do in the lab test mode so ID in the middle of the lab which is trees almost three hours you're gonna have a break and the computers are gonna get into lab set test lab mode in the test lab mode you don't have access to internet or communications with others cell phones is this strictly prohibited so please make sure you you read and learn about the rules because TNA's are instructed to give you zero points if they see you using your cell phone just treat your lap kiss mode as your actual exams so lab exams so the first half of the the lab you're going to walk through with that pre lab PDF that I'll post before the lab so you're gonna start using the Vrba's simulator and so let me just so this is the this is a simulator so these are some assembly instructions that you're gonna learn start to learn as of Wednesday so when you when you write some assembly anybody Spy is kind of compiled it's like an actual simulator 4 X 5 so say in your instructions you were given to write and add AI codes right you can just write it down here on the first pre lab and then compile it and see the output and then when we came back from the break in the middle of the lab session so after 90 minutes the second 90 minutes which is sort of alias the second part of the lab your computers will turn into the test lab mode so you're going to cut access from the outside board and you're expected to follow through the rest of that PDF which are the actual questions and then at the end of the lab you submit a file a text file of the output of is you just go to the files to the same as save it as a file and submit it in the lab I'll give you the instructions there there are also available on the website so just spend five minutes you're gonna learn everything where to ta is available there I myself would be there as well so hopefully we don't run 20 issues questions I don't think so I mean I don't mind but the point is there are three sections in this course and around 400 students we want to make sure the fairness is for everyone otherwise if you haven't notes or if you don't it doesn't change my life you know since we decided that you can call accents from the outside and you've got you bet the instruction the pre-lab a few days before perhaps and inside the pre-lab so you should be able to start you know testing yourself in the test and it's just laughable so 99% know if something happened now I'll let you know in the last you know it's just not it's not a matter of taking notes or having a calculator or having a scientific calculator or having a calculator on your cell phone it's just some ground rules that we've set and then we follow through on that I don't think so I mean if it's really needed I can set up Moodle but I think Ricky on Vicky I can set another tab and then I just send you by email or make it available a certain way you can access it yourself me but if Moodle has a specific functionality that you definitely need to use it let me know email me and I'll consider no worries you'll find out what a great yeah that's the last thing you have to you know worry about no worries so yeah no not these Peaks but it's starting from next week you can have your first lap so a couple of days before I'll I'll update the pre-lab for you guys you can also download this and ice it from your computer at home or you even connect to the lab machines there are several ways you can start working on the and using the manual on that you're actually physically there yeah no I mean in the test lab mode and also in the pretty labs you need to be physically there yeah okay all good cool all right so we started last week yeah actually yes and no because if something goes wrong if your file is corrupted if I don't know an earthquake happens and the electricity runs out and whatever we can have as sort of a follow through of your relapse and and see how you were progressing so that's a good trademark for us to see your progress so that's why you need to submit the pre labs as well so that's that's mandatory your evaluation would be based on your test materials but given that we have the pre lab material as well so you need to have that one as well okay all right so for the office hours I'm obtaining an office and then I'll let you know the exact time by Wednesday and also the TAS we have assigned the TAS but for a specific hours just check again on Wednesday I'll update the website as well and also I'll take the pizza slice and I'll put it back as a revised version on the the wiki page as well so again just make sure if you're only enthuse it so that was a reference we started to have a look this will change so we reach up to the point that we were we were talking about opening the box inside a computer and see how we gonna speed up their processes the tasks and the instructions running on the computer and we're on we were trying to understand and sit said stone about some different methods to understand performance in algorithmic level in programming language or compilation level in processor level input and output or operating system-level light so release off to the point that we were talking about an assembly using a compiler and assembler and how does the assembler after that it's gonna generate machine read you know binary code and then we actually run that on our hardware right so we are trying to familiarize with this process in this course we were talking about a big picture of you know components of a computer as an input/output control data path and then we carried on with the the high level view of starting from transistors and then how they derive from a silicon chip that yells performance for your computer so all right so this session began it you're gonna carry on this introductory conversation and we hope we can finish the the first chapter today and on Wednesday we're gonna start the first chapter which is the the beginning of instruction set architecture of risk file and you're gonna learn gradually how does the assembly you know working increase five actually alright so if you recall that video that I showed at the end of the previous lecture so from the from the sand we extracted the silicon and we made an alloy with that and then we slice it and whatnot so we're up to the point that we reach to wafer so i7 core with this 32 nanometer technology right so the companies that fabricate and produce this wafer needs to have some formulas in order to understand what's going to be the output they're in our product so let's have a look at three simple formulas regarding that alright the first one is pretty straightforward so the cost hair dye is actually your cost per wafer / dyes per wafer and they yelled which is the the amount of dyes inside and wafer that they're actually functioning on right coming back to that figure again so it's starting from the silicon ingot they they made an alloy they slice it they design the the blank wafers and they repeat this process several times and they make a pattern out of that so when it comes to test that pattern right so they need to repeat the process with the Dicer and then test the dies again so this process we can define a yelled metric that understand the proportion of working dyes parameter say if one of them doesn't you know anything so you're out of six your five out of six yellow from the performance you need a chair yeah there's another one here as well you can take this why should there's another one here all right so having having to yell mad free you can compute the cost per die right it's gonna because per wafer divided by cost per wafer multiply the yelled of that dice per wafer so the second formula approximates almost the dies per wafer is approximately equal to wafer area divided by diarrhea does anyone know why you do not have proximation and it's not an equal sign because it makes sense you should be equal right because I'll wait for our circular right so what happens in the corners in the corners we miss almost a little bit of our area so that's why he's almost as an approximation for that and a third one is coming actually from the empirical results of companies when they produce the chips so it's actually an empirical data on integrated circuit factories and that the the exponent related to the number of critical processing step right and as you see it's not linear so now at least we have very preliminary formulas in order to understand how effective we've been producing dyes and you know wafers in into our integrated circuits right all right now let's go back for another example so we were talking about defining performance in the previous lecture so now consider we have four different planes so of them all like bees I believe he was in his 70s and 80s it was for British perhaps air and sea perhaps stands for operation or something you see yeah so it was a British aircraft in the seventies and so we have two different wings 777 and 747 so you see regarding different metrics so passenger cap capacity the range the cruising range so how far can the plane go without refueling and cruising speed and also the passenger mile same right so it just defines in what way we are looking at defining performance for our ICS and for our CPUs for our computers right how fast that process can can run how big of a how big of an CPU we want how much money we want to spend how much processes should have been done in order to make that right so this is small analogy lead us to define some more you know interesting metrics so first of all if you're talking about computer architecture and organization and programs so for us definitely response time is of importance so what is actually respond so it defines how long a program takes to run its task or in a way how long it takes to do a task on a server computer within program right so the other one that is very important is the true put of that hardware so the total work done per unit time right so we are interested in completing certain amount of workload in a certain amount of time right so it gives us a perfect metric to understand that as well so for instance how many tasks how many transactions how many instructions where you know processed input output per hour or per certain amount of time right so given these two we can define questions such as how our response time and throughput affected by other you know other variables so if you replace a processor with a faster version say they replace itry with five still so what are we expecting to get with the newer processor what if you change from i-5 to i7 r i9 right how about we add reports we start from a single core to a multi-core or a mini core all right so let's focus just for now let's focus on the response time for now okay so in in a matter of performance an execution time when we say when we say X is n times faster than Y right this actually means that the performance of x divided by performance of Y is actually n right so the higher the performance the faster it is so you see that he's having a relation the faster the performance is the lower the execution time so for instance time taking to run a program it's gonna take 10 seconds on a it's gonna take 15 seconds on B right so the execution time of B divided by execution time of a it's gonna be 1.5 sauce if we you know point of view we're gonna see that 1.5 a faster than B right is it clear all right so now that we know the relation between performance and execution time let's talk about some other more detailed measurements for execution time so we have elapsed time which accounts for total response time including all aspects from the point that you input the point that it was processed in the OS the all the overheads of the OS and then the idle time if the OS was putting the program in to wait or halt or some some cycles of pause and then after the point that to receive the output right so it's gonna be a elapsed time there is another one inside elapsed time that accounts only for the CPU time right so some portion of that elapsed time might account for the CPU so and that's exactly referring to time spent processing the job only and other non processing factors from that elapsed time all right so let's focus on the CPU time for now so this figure shows that depending on the clock of your CPU when you say you have a CPU of 4 gigahertz 2 gigahertz or in the past it was in a matter of megahertz so it shows how fast the clock cycle of the CPU work right so the faster the clock cycle the smaller these fluctuations write z1 ones so each time it goes 0 and 1 so we're gonna call it one clock period right the faster the clock is the smaller or your clock period so faster they can have cycles so as an example when you have a clock period of 250 heterocycle I'm sorry para second it's as if or 250 multiply I'm sorry it was Pico and not para because it's - yeah it's 250 12 seconds right so now if you want to calculate the clock frequency or the rate of the clock so actually 4 gigahertz now you know that in the stands for because of 9 so you actually understand that at each second you have 10 to the raise of 9 times clock period for your CPU right so given this given this a small introduction we can define CPU time as a form of clock cycles and number of clocks it takes for you know for a certain task to complete its CPU time right so CP Tom stands for CPU clock cycle some some Hertz multiply clock cycle time each of those or you can define our CPU clock cycles divided by the rate of the clock right so the higher the rate the smaller D the CPU time if you change your CPU from I don't know 1.5 to 2 you're expecting to have some performance increase by how much we're gonna we're gonna have an example shortly and let's see if we if we you know change our CPU say without a new CPU and a cheap you know speed is from I don't know in the past behind - you got hurts now it's 4 gigahertz right are we expecting it to week 6 speed up or not let's see how does the relation between D the clock rate and the overall signal kinda goes in computer architecture any questions so far right all right so example we haven't we had a computer it used to work at 2 gigahertz so as a task or a program was taking 10 seconds right so his CPU time was 10 seconds now we've got a designer he wants to design a new computer a new chip for us right and he's trying to aim for 6 seconds as the new CPU time for the new task so we are interested in you know having our task finished by by the six-second instead of the temp right so he's trying to design the new chip and he's coming up back to you and says ok I can do the faster Clark but this causes one point to multiply the clock cycles for the rest of the CPU design right I need to increase that per day but that amount in order to you know include certain more functionalities in order to do that job in 6 seconds so now let's see the clock rate of the new computer now so the clock rate as I talked in a previous slide it's going to be Clark cycles of the computer be the new computer divided by the CPU of the new computer we were aiming at six seconds the new clock cycle is one point two of the previous one right and then by the time we calculate the clock rate of the new computer we see that in order to offset this time and make it faster from ten seconds to six seconds given the other constraints no works at four gigahertz so we are doubling the processors but not we are not you know having the CPU time in this case so you see that it's not that easy just to double the amount of CPU double the speed of the CPU of double the clock speed of the CPU and the next big everything is going to go double speed now right there are other constraints involved and we're gonna learn more and more about that in this course is it intuitive okay now let's add one more level of detail and this performance evaluation so we talked about the the task let's zoom into task and define what tasks are actually mean so how do we define tasks are they lines are they are those cycles are there instructions how do we define those so one of the better metrics in order to define those is instruction count so the number of instructions your compiler outputs from a high-level language so you write printf hello world so how many instructions your compiler with out put that into assembly code right so we actually count those and it stands for instruction count or IC so a better way to define clock cycles is actually by the the multiplication of instruction counts and cycles parents so how many instructions I have and how many cycles it takes per instructions you know to to process that so you multiply these do you have your sign now go back to the previous formula for the CPU time now we can you know just add this into the into the into its position and then we're gonna have CPU time is equal to instruction count CPI which is the instruction cycle per instruction so cycle per instruction multiplied clock cycle time and then actually it is equal to instruction count multiplied clock per instruction total amount of work and then divided by the clock rate the faster the clock rate the lower the final value would be which is the clock side which is the CPU time right so if you increase your heart rate process faster so now we have different components to play with if you want a faster instructions account so we need to have a better I say a better compiler that outputs more efficient codes lower number of fewer number of instructions if we and then the average cycle per instruction is actually completely determined by the the architecture of the CPU how fast is gonna process that right does it use the pipelining does it use other optimization does it use forwarding or or anything else that we're gonna learn later on question oh yeah okay so let's see an example about this there pretty intuitive and we just have to learn it by by feeling out say you don't have to memorize anything it just you need to compute all the world workload and then speed you want to process that you know and that workload could be you know into more fine-grain and into more details can I expand that into different layers all right so computer a it has a cycle time of 250 PS these cycles per instruction is to write and computer B has a cycle time of 500 PS and a CPI is one point to write both having the same Aiza or is a that means that their architecture for instance both of them are having risk 5 I I say they have the same you know 64-bit for instance I say so let's see which is faster in order to compute that we need to compute the CPU time for a and CPU time for be so CPU time as you recall from the previous slide it was the modification of instruction count for the put the overall instruction which we actually we don't need it because it's gonna it's gonna be both for the same and it's gonna and you know we don't need the number actually so the thing that is important for us is only CPI a and a cycle time for the computer a so I stands for that so it's kind of take to multiply 250 and then on the case of the B is kinda take 1 point to divide and multiply 500 right so now you see that computer a is actually faster but looking at the first one we didn't sound like right and then if you divide the CPU times of B by a you're gonna see that by how much it is faster so 600 divided by 500 it's one point to fast right so that's just for a single as a single application or a single instruction counts what if you have many of those and each of those have different CPI questions instruction set architecture such as risk yeah yeah or meap's or arm yeah okay good all right so now let's talk about what if you have many castes right and each of them has his own CPI so if we go if you zoom in again into into that if you have different instruction classes that each of them each of them takes different number of cycles so you learn in chapter two that normally load instructions are taking more time than tape controlling instruction is taking less time than load for instance in general and it depends on the architecture you see different instructions take different amount of cycle in CPU to to process so what if we want to have them in more fine-grain right in this case you have to sum them up and weight them by their own CPI right so each instruction count of I it has its own CPI I and then in order to find the overall sum from I to n and weigh them with their own CPI and at the end when we want to calculate the final CPI it's gonna be the weighted average CPI is it clear alright let's see an example of this so we have now we have two classes right so a CPI row this row is provided by the hardware designers so as if you have two different hardware's a B and C right tree actually and then the instruction count the number of instructions are output by the compiler you have so the compiler has to out emit code instructions when you know compiling a high-level language like C C++ Java or anything else right now let's see by these two dimensional data two different Hardware different compiler right when we have two different sequences of I see two different instruction programs which one is going to be faster right in each class a B and C so for the sequence 1 we have instruction count equal to 5 y is 5 yeah and for sequence to the instruction count is 6 right and then you have to make a weighted sum here so to multiply 1 1 multiplied 2 to multiply 3 is equal to 10 so the average CPI the fine write on a sequence number to a different compiler for instance it would be 4 multiply 1 1 multiply 2 & 1 multiply 3 Sony the answer was 9 and then it would be the average CPI is 9 divided by 6 okay you see that when we wait them depending on the amount of cycle each test takes you have different output now right okay so now let's put everything together in a single I would say more beautiful formula so CPU time the overall picture CPU time stands for instructions / program multiplied clock cycles divided by instructions multiplied seconds divided by clock cycle you see the more we go right the more fine-grain we go right choice starting from instructions per program clock cycles per instructions and then seconds + / clock cycle right so now you see that the performance of a computer depends on many things possibly can affect the CPI programming language definitely affects IC right compilers for sure effects I see depending how long how big of a code they're gonna generate how many instructions they're gonna generate out of a single save print right and then the instruction set architecture or the is a affects the I basically CBI I see time as well okay does this make sense to you I'll pause here for some seconds okay so now we talked about we talked about the big picture of CPU time so let's see how much of power is gonna consume when you're targeting certain performance right because power is of an important issue recently if you if you're a savvy computer scientist or an engineer you know that we are in a second power wall sort of era and power trends are showing us that you're actually lowering the power because we can't afford to have you know higher performance you know chips within the same area that they they take in order to perform that if you want to make this smaller and smaller we can't afford to just you know going hiring the frequency at the same time we need to make them low-power as well otherwise the area becomes larger and larger so the big picture of this definition is called power wall and starting from the mid 2000s they face the first power wall so I'll talk about it later but for now you just see the trend here in 80s they were just going up in a clock rate the blue line and at the same time the power would go followed almost the same trend up to the point that paint Pentium 4 I'm not sure if you recall it is for my age yeah when I was a was perhaps primary school perhaps yeah that was that was the best iced I nine we had so and Pentium 4 was was a benchmark for speed at a time so you see after pinion 4 when the clock cycle reached three three point six gigahertz it was pretty much very high at the time you see that now the best views are are no more than four because they can't afford to pay the price for the power so you see them the the clock rate actually went down because that was the first time they hit the power wall Pentium 4 was a very performant but it was consuming a lot of power and they couldn't simply afford to go higher because the power would just surpass it no 100 watts it was a lot you can't afford to have a computer a personal computer that just you know with a single cord to spend this amount of power so after that you see that they they try to make the process more optimized you got a little power they define new your generations of you know you know family of chips the one that are look more low power are using in your cell phone devices and many-to-many some of the laptops and the more high I perform and wanting the servers and some some personal computers as well so so you see that the power consumption is dropping and they almost keep the same rate right so let's see what was the reason behind this so roughly speaking consumption of power is related to capacitive load of the material with transistors multiplied voltage to the raise of two and multiplied frequency so in CMOS technology so CMOS is a technology underneath the your IC it stands for complementary metal-oxide-semiconductor right so it's four they're using that technology for constructing ICS including any ICS like microprocessors or anything else right so the dominant technology for ICS is a CMOS actually and for CMOS the primary source of energy the the primary source of energy consumption for CMOS is actually coming from dynamic energy so energy is the energy takes the transistor to turn itself from 0 to 1 or 1 to 0 right all we need is material which we found finally in transistor when we switch it 0 1 and 1 to 0 it's kind of you don't perform the task force because all the codes that are emitted by assembler or 0 and 1 right we just have to mimic to zero and ones with a material and that material consumes power you have to make that process more efficient right that's the whole I'd say story behind computers right you have all these zero and ones try to find it here so this is the actual program that computer runs right zero and one so you need something that performs these switches zero and one and by each switch is gonna consume some energy right and using that you can run things on computers so alright so I was talking about the the major source of power for CMOS and that is called the dynamic energy that is energy that is consumed when transistor switches their state from zero to one and vice versa right and that depends on the capacitive load of that transistor with the frequency and the voltage applied so that's just estimation so as an example if you want to have power to increase by 30 weeks right we need to have our frequency to increase by a thousand times at the same time we need to make that voltage consumption from five to one right this is a pretty hard task normally and that's what it takes in every generation of computers they want to make it more power efficient your laptops perhaps I don't know uses less than hundred Watts 70 watts even less yeah the the pieces they did the newer GPUs of Nvidia around hundred 150 watts depending on the workload so because they can't afford to spend more power on that even the newer version of CPUs having multicores and many cores there is another problem they called dark silicon so let me see if I can void something here so say this is your eight core so this is your eight core CPU right core zero core one two three up to seven so you have spent some money that's a good CPU you got five six right so in most of the architectures there is a well-known problem I mean on all the architectures we can't afford in most of the cases depending on the city or governor we can afford to turn them on at the same time all the time first of all their heat that their generate affects the performance of the adjacent course right and if they are at 100% load the power consumption would be very very high so the battery issues the power issues and so on and so forth so the governor's you can see even in your CPU depending on the governor on your I'm sorry your cell phone if you install some applications to monitor your CPU governor for instance cpu-z CPU I see you can see that not all of your cores are on at the same time right so perhaps these terrorists off is on this is off this is off this is on and then they try to make the adjacent one not on at the same time because they have different you know impact on onion adjacent so this problem is known as dark silicon and there are many researches going on researchers have been trying to address this so this is just one of the issues in in multi-core domain so let's go back to our you know it's your our course which is still in the single core up to here so you see that by increasing the power we need to go a long road in order to optimize that process is it clear okay so suppose your new CPU has 85 percent of the capacitive load of the old CPU and the designer is managed to reduce 15 percent frequency reduction right so the power of the new after all these optimizations has 0.5 to of the previous generation right and that was a powerful I was talking about some minutes ago so after the Moore's law when when I mean we are at the at the error that actually Moore's law is now is no longer an issue for computer designers because we can't afford to turn them on at first we have hit the the power wall so optimizing the the voltage frequency in capacitive is of an you know more important issue than just increasing the clock rate right so you can see it actually here so in the good old 70s and 80s you see on average 50% increase right after hitting the first power wall that increased I've gone way way slower to around 20 percent right and it's just this graph only shows the uniprocessors we're not talking about the the multi-core or many course by now we'll touch on that later right sort of 52 percent where the clock speed is an increasing that is over now right as I mention and it's now more architectural innovations multi-core many cores so if you just think about like 30 years ago people will just write software's and they will just wait for a new CPU to come out yeah let's wait for in two years our software will become faster on its own right but now we have to optimize it in different levels you have to optimize the compiler you have to optimize the way we write code your algorithms new software paradigms are using OpenMP MPI are we going to paralyze it and other issues all right that lead us to a multiprocessor area which is the the the current era right so every chip has more than a single processor so why this is that leaders are very good to start but this raises even more issues for us now how we're going to define tasks for those how we're going to define the the constraint between the programs do we have dependencies in our failure for loop you have a for loop you want to do something in a for loop some variables assignments and so are we having dependencies between the assignment of the variables are we running into issues of that are we going to read after a right right are we going to load some variables that has been assigned in another tread at the same time and you're losing the communication so these yells another set of issues for itself and perhaps you're gonna learn it in I'm not sure if you have a pearler programming course or something do you have something yeah so that's going to be an interesting one so for that you need to load balance you need to define new software paradigms programming for performance and optimizing the to do the do the tasks right okay I have a few more slides I'm gonna just wrap it up quickly so spec is an organization you can just find in the website ladies spec orgy they are trying to standardize benchmarking of computers and you have different benchmarks so for CPU they have is one of the latest one is spec CPU 2006 they have benchmarks for integer variables and for 14-point variables as well so the way they do is trying to normalize away different tasks or different programs different standard programs as benchmarks are running within a certain architecture and they report the number by averaging them but as you know when you have a speed ups you can just simply average them by arithmetic mean you just can't multiply them and divided by the number there the correct way to average speed up is by geometric mean so that that's the formula for a geometric mean if you use our know MATLAB Python or Excel or whatever so when you're calling mean there is geo mean as well there is harmonic mean as well so Deb another application so but in general as a rule of thumb when you have a speed-up numbers one point two one point seven whatever when you want to try to compare different a speed-up depending on from which angle you look at the speed-up so say if you have two computers right if you base your speed up on the first computer or the first or the second computer when you average them try to do the math and you see that you have different numbers so that's why on the case of averaging a speed up numbers you need to use geometric mean and that's the formula for that so spec uses that to normalize really tip to reference machine and they have reported for the integer values and also they have another one for the floating point values so these are different standardized softwares right a video compression software parsing so these are different benchmarks that they run on each of the architectures in order to capture their performance right and they make an overall average and they call that overall workload overall ssj underscore ops and when they add the power metric over that so the most accurate way we have for benchmarking CPUs is overall ssj underscore our third what so how much it takes paraquat right and then each of those comes from its own averaging so using that they publish product and pay your weekly for each of the CPUs that comes out I seven nine nine nine five or different as well forearm and you can just find it on the website so this is one way to benchmark the performance right it's actually simplify the marketing of computers so when you wanna buy a computer you know actually how fast is it because if you want to buy a computer and just ask the guy I want a fast computer he's gonna ask how you define your fast right you want a low power computer you want a fast CPU cycle computer know how do you define that test alright so the last point we're gonna talk about today before wrapping up the chapter one is the famous and as long right it's very intuitive perhaps you've already considered sometimes when you were thinking about optimization test hopefully so when you do have when you try to have an improvement over a task right you always need to take into account how much of the overall workload has been improved by the new optimizations and I write say you have a huge chunk of I don't know this is this just bottle of water right if you're improving half of the water if you found a way to you know devise a new material called in your water that takes half of the weight of a water right and you change half of this bottle of water with that new material that has half of the weight of a water right what's going to be the overall improvement you had simply you have to take into account the amount of improvement there overall you know material right and that's the the premise of and as well when you have an improvement it doesn't affect all of your workload right if 10% of your workouts has been doubled it doesn't mean that your hundred percent energy level is that ten percent so you take a fraction of that and then you add the tea of unaffected one the unoptimized one right is that clear all good so that's the that's the simplified version of and as a lot they have defined that many different version of that for many core multi-core and different domains so say yeah and as I talked before it depends there's another method for you know your workload is that how much of performance of your CP or you using are using hundred percent of your CPU are using fifty percent load of your CPU who are using at ten percent right so this is another metric that comes into play in order to compute the performance of your machine finally there's a well-known metric called mix and knead system for millions of instructions per second so for a long long time computers very different is a s where benchmark by the number of MIPS there were outputting right so the number of us watching counts an execution was for millions of instructions so that's that's why you have ten to the raise of six number there right okay so by now as an introduction we learned the cost and performance as a trade-off of improving CPU right we just have to take into account the all the file all the fine details changes we doing hardware and software as well and that's the real premise of you know computer architecture and currently power is a limiting factor right so we use different are parallels and improve our performance all right any questions nope next pitch yeah so on Wednesday are we gonna start chapter two and by the end of the week all closely the pre labs for you next week laughs next week
**Course Overview and Lab Introduction** - Welcome and semester start - Course materials uploaded to wiki page - Section E 221 P details - Lab registration and attendance instructions - Lab schedule: first lab next Monday, pre-labs released before each lab - Use of simulators: first 4 labs with RVs simulator, post-midterm labs with Verilog simulator - Teaching assistants (TAs) assigned per lab section - Lab rules: test lab mode disables internet and phones, strict enforcement - Lab process: pre-lab instructions, simulator exercises, then test mode with output submission **Computer Architecture and Performance Concepts** - Introduction to computer components and instruction set architecture (ISA) - Silicon wafer fabrication and yield formulas: - Cost per die = cost per wafer / (dies per wafer × yield) - Dies per wafer approximated by wafer area divided by die area (circular wafer approximation) - Empirical yield formula depending on processing steps **Performance Metrics and Evaluation** - Response time: time to complete a task - Throughput: total work done per unit time - Performance inversely related to execution time: $\text{Performance} = \frac{1}{\text{Execution Time}}$ - CPU time components: - Elapsed time (total response time) - CPU time (time spent processing only) - CPU clock and frequency explained: faster clock means smaller clock period - CPU time formula: $$\text{CPU Time} = \text{Instruction Count} \times \text{CPI} \times \text{Clock Cycle Time} = \frac{\text{Instruction Count} \times \text{CPI}}{\text{Clock Rate}}$$ - Example comparing two computers with different cycle times and CPI - Weighted CPI calculation when multiple instruction classes exist **Power Consumption and Limitations** - Power wall concept: limits on clock speed due to power and heat - Power consumption formula in CMOS technology: $$P \propto C \times V^2 \times f$$ where $C$ is capacitive load, $V$ voltage, $f$ frequency - Dynamic energy consumption when transistors switch states - Dark silicon problem: multi-core CPUs cannot have all cores active simultaneously due to power and heat - Trends showing power efficiency improvements but clock speed growth slowing **Multi-core and Parallelism Challenges** - Transition from single-core to multi-core architectures - Software challenges: dependencies, parallel programming, load balancing - Importance of new programming models (OpenMP, MPI) **Benchmarking and Standardization** - SPEC organization benchmarks for CPU performance (SPEC CPU 2006) - Use of geometric mean to average speed-up metrics - Benchmarks include integer and floating point workloads - Overall performance metric SSJ_ops for benchmarking **Amdahl's Law for Speedup** - Overall speedup depends on fraction of workload improved - Formula: $$\text{Speedup} = \frac{1}{(1 - f) + \frac{f}{s}}$$ where $f$ is fraction improved, $s$ speedup factor - Highlights importance of optimizing frequently executed parts **Summary** - Course covers computer architecture basics, performance metrics, and practical lab work - Emphasis on balancing speed, power, and complexity - Upcoming chapters will cover ISA and assembly language in depth If you want to access the full transcript press the button bellow