Download PDF - IBM Redbooks

Transcript
The reason we have to “back off” the addresses prior to going into the loop is because the
lfdu and stfdu instructions generate the effective address first by adding the displacement
value (8) to the contents of the GPR. The result of this addition is placed back into the GPR
(for example, rA = rA + 8). If not, the loop would skip over the first 8 bytes (64 bits) of the
matrix. The fmadd instruction, first introduced in the POWER architecture, performs a multiply
and add operation within one instruction. Because normalization and rounding occurs after
the completion of both operations, the rounding error is effectively cut in half as compared to
doing these as separate instructions (for example, multiply instruction followed by an add
instruction).
1.7 Superscalar and pipelining
To this point, the paper has described registers and instructions. One major feature of the
PowerPC microprocessors is to execute instructions in parallel. Often the term superscalar
appears in the description of an IBM Eserver system and are not quite sure what that term
represents. For those readers who are actively or considering writing assembly language
programs for the PowerPC 970, a brief description is presented here.
Figure 1-1 on page 8 showed the block diagram of the PowerPC 970 with its 10 pipelines
(CR, BR, FP1, FP2, VPERM, VALU, FX1, FX2, LS1 and LS2) for instruction execution. The
PowerPC architecture requires a sequential execution model in which each instruction
appears to complete before the next instruction starts from the perspective of the
programmer. Because only the appearance of sequential execution is required,
implementations are free to process instructions using any technique so long as the
programmer can observe only sequential execution. Figure 1-6 on page 23 shows a series of
progressively more complex processor implementations.
22
IBM Eserver BladeCenter JS20 PowerPC 970 Programming Environment