Download PDF - IBM Redbooks
Transcript
The reason we have to “back off” the addresses prior to going into the loop is because the lfdu and stfdu instructions generate the effective address first by adding the displacement value (8) to the contents of the GPR. The result of this addition is placed back into the GPR (for example, rA = rA + 8). If not, the loop would skip over the first 8 bytes (64 bits) of the matrix. The fmadd instruction, first introduced in the POWER architecture, performs a multiply and add operation within one instruction. Because normalization and rounding occurs after the completion of both operations, the rounding error is effectively cut in half as compared to doing these as separate instructions (for example, multiply instruction followed by an add instruction). 1.7 Superscalar and pipelining To this point, the paper has described registers and instructions. One major feature of the PowerPC microprocessors is to execute instructions in parallel. Often the term superscalar appears in the description of an IBM Eserver system and are not quite sure what that term represents. For those readers who are actively or considering writing assembly language programs for the PowerPC 970, a brief description is presented here. Figure 1-1 on page 8 showed the block diagram of the PowerPC 970 with its 10 pipelines (CR, BR, FP1, FP2, VPERM, VALU, FX1, FX2, LS1 and LS2) for instruction execution. The PowerPC architecture requires a sequential execution model in which each instruction appears to complete before the next instruction starts from the perspective of the programmer. Because only the appearance of sequential execution is required, implementations are free to process instructions using any technique so long as the programmer can observe only sequential execution. Figure 1-6 on page 23 shows a series of progressively more complex processor implementations. 22 IBM Eserver BladeCenter JS20 PowerPC 970 Programming Environment