Download 1130-CA-06X - about the IBM 1130 Computing System

Transcript
IMIIMME■N.
roc
■■•
■MIEW
MM■
■IMIMINIM
Application Program
1=1
H20-0333-1
1130 Statistical System (1130-CA-06X)
User's Manual
The 1130 Statistical System performs four major statistical
functions — regression analysis, factor analysis, analysis of
variance, and polynomial fitting.
This manual contains, for each type of analysis performed,
a description of the computational algorithms used, the form
and content of the control cards, operating instructions, and
sample problems.
Kristofer Sweger
I 1111wci
nn
11..1111
11111
NIFFI
111111111111
Second Edition
H20-0333-1 is a major revision, obsoleting H20-0333-0 and incorporating
TNL N20-1029-0.
Specifications contained herein are subject to change from time to time. Any such
change will be reported in subsequent revisions or Technical Newsletters.
Copies of this and other IBM publications can be obtained through IBM branch
offices. Address comments concerning the contents of this publication to
IBM, Technical Publications Department, 112 East Post Road, White Plains, N.Y. 10601
© International Business Machines Corporation 1967
INTRODUCTION
The 1130 Statistical System contains four major analysis programs:
1. Stepwise Linear Regression
2. Factor Analysis
3. Analysis of Variance
4. Polynomial Fitting with Orthogonal Polynomials
Each of these analysis programs is composed of a number of subroutines, which are
stored on the disk and are called into core storage when required for the execution of a
particular job. The logic flow of the programs and the type of analysis to be performed
is controlled by a main program, which reads the user-supplied parameter cards and
calls in the appropriate link at the proper time.
Although the programs imply different techniques, a common approach can be used in
executing a job for any analysis. Chapter 2 of this manual is divided into four parts
that describe completely the necessary parameters and monitor control cards for
execution.
Special features available with this package provide added user flexibility:
In all four major programs, data card formats can be specified by the user.
•
Stepwise Linear Regression. Matrix input and output are allowed with a pooling
option which provides for the combining of raw cross products matrices, either by
addition or subtraction. This allows the combining of input from different sources,
or of input that is available at different times, without requiring recalculation of
these matrices. The subtraction feature gives flexibility for the handling of outlyers.
Residuals are available on option.
•
Factor Analysis. The pooling options, and matrix input/output, are available, as
described above, for stepwise regression.
Factor scores are calculated on option, and punched on option.
Several options for the handling of communalities are available.
Oblique and orthogonal rotations are allowed.
• Analysis of Variance. A table generation feature allows output from the factorial
design analysis in standard format.
• Orthogonal Polynomials. These can be calculated for both equally spaced and
unequally spaced intervals.
Derivatives are calculated on option.
Scaling of input is allowed on option.
CONTENTS
CHAPTER 1: GENERAL OPERATING INSTRUCTIONS
11
System Generation
1.2
Control Cards
1.3
Program Pauses and Messages
1.4
Stacking: Sequential Program Operation
1.5
Typewriter and Punched Card Output
1.6
Machine and Systems Configuration
1.7
Programming Language
1.8
Reference Material
1
1
1
4
4
5
5
5
6
CHAPTER 2: PROGRAMS
2.1
Stepwise Linear Regression
2.1.1 Summary of Output Statistics
2.1.2 Job Execution
2.1.3 Data Input
2.1.4 Matrix Input/Output
2.1.5 Operating Instructions
2.1.6 Sample Problem
2.2
Principal Components and Factor Analysis
2.2.1 Mathematics of Principal Components and Factor Analysis
2.2.2 The Program
2.2.3 Summary of Output Statistics
2.2.4 Job Execution
2.2.5 Data Input
2.2.6 Matrix Input/Output
2.2.7 Operating Instructions
2.2.8 Sample Problem
2.2.9 References
2.3
Analysis of Variance
2.3.1 Tests of Significance
2.3.2 Job Execution
2.3.3 Analysis of Variance Table Generation
2.3.4 Data Input
2.3.5 Operating Instructions
2.3.6 Sample Problem
Least-Squares Curve Fitting by Orthogonal Polynomials
2.4
2.4.1 Summary of Output
2.4.2 Job Execution
2.4.3 Data Input
2.4.4 Operating Instructions
2.4.5 Sample Problem
General Notes on the Programs
2.5
2.5.1 Transformations
2.5.2 Notes on Correlation and Eigen Analysis
2.5.3 Punched Matrix Output
2.5.4 Scaling
7
7
7
8
16
17
19
21
31
35
42
45
46
56
57
58
60
71
72
76
78
81
83
84
86
89
92
93
99
100
102
111
111
111
112
113
CHAPTER 3: GENERAL FLOWCHARTS
CHAPTER 4: SAMPLE PROBLEM TIMING
CHAPTER 5: ERROR MESSAGES
114
117
118
CHAPTER 1: GENERAL OPERATING INSTRUCTIONS
1. 1 SYSTEM GENERATION
After preparing a disk with the monitor system (1130 Disk Monitor System, 1130-0S-001,
described in manual C26-3750), the distributed source cards, preceded by a cold start
card, can be loaded into the card reader hopper, and compiled and stored on the disk.
It is not, however, necessary to store the total 1130 Statistical System on the disk
before using any one of its four major programs. Within the discussion given for each
specific program is information pertaining to loading the particular set of routines
necessary for that analysis type.
The distributed decks consider that the operating system will use the 1132 Printer as
output. If the console typewriter is to be used as the output device, monitor generation
must consider this, and the IOCS card in each main program must be replaced by an
IOCS card stating:
*IOCS (CARL), TYPEWRITER, DISK)
Identifying information for this exchange, which should take place before loading the
statistical system, is listed below.
Program Name
COREL
POLY
POL2
REGR
REGR2
ANOVA
ANOV2
FCTR
FCTR1
FCTR2
FCTR3
IOCS Card Indentification (cc 73-80)
CORL 10
POLY 20
POL2 20
REGR 20
RGR2 10
NOVA 20
NOV2 20
FCTR 10
FCT1 20
FCT2 20
FCT3 20
In the routine PRNTB
a. Card PRNB 150 should be changed to read LIBF TYPEZ.
b. Cards PRNB 70 - PRNB 130 should be omitted.
In the decks distributed with this system, identifying labels are given in cc 73-76. These
four characters do not allow labels to agree perfectly with names of programs. When
referencing programs, keep this distinction in mind.
1.2 CONTROL CARDS
Each of the four programs included in the 1130 Statistical System requires monitor and
program control cards. These cards are described in the job execution section of the
specific program being considered. However, certain program control cards are
standard for all programs. Their descriptions follow.
1
For any control card, numbers specified as integers (I) (this includes all numbers used
for program control) should be specified as follows:
numbers should be shifted to the right of their fields (right-justified) unless
left justification is specifically called for.
1. All
2. Blanks and zeros are synonymous.
STANDARD PROGRAM CONTROL CARDS
Input/Output Units Card
The function of the input/output units card (Figure 1) is to assign logical unit numbers
to each of the I/O devices used throughout the program. Each subroutine that requires
the use of an I/O device has been programmed with symbolic unit designations. This
card fixes a number to a specific I/O device.
Column
Meaning
1-2
The unit for input of all control cards and source data.
Normally, it is set equal to the logical number of the
1442 card reader, which is 02.
3-4
The unit used for card output of computed matrices. Normally,
it is set equal to the logical number of the 1442 card punch,
which is 02.
5-6
Output switch
0 - 1132 Printer output
1 - Typewriter output
CC: 1 2
102
34
02
56
00
Figure 1. I/O units card (printer)
Job-Title Card
The job-title card (Figure 2) allows the user to assign a job number and title information
for the job to be processed. This information is used only for labeling and is not used
for processing in any program. The job number and title contained on the card are
printed as the heading line on each page of output produced. The job number appears
in the first four columns of any punched card output produced.
2
4
CC: 1
21
9
Multiple regression for class 3 data 10/7/64
Figure 2. Typical job-title card
Column
Meaning
1-4
Job number
5-8
This field is not used.
9-80
Title information. These columns may contain any legitimate
key-punchable characters that serve to identify the job.
Variable Format Card
Each program was designed to allow some flexibility in the input of data. Although
1130 FORTRAN does not allow the use of an object time format definition, a specially
written format processing program is employed to enable the user to specify the format
of his data by means of a FORTRAN-like statement. The format statement may contain
almost all the specifications included in a normal FORTRAN format statement (as
described in 1130 FORTRAN Language (C28-5933), pp. 11-15), with the following
exceptions:
1. Only I, E, F, or X data specifications are allowed.
2. Continuation cards are not allowed.
3. The use of a slash (/) is not permitted.
4. Internal parentheses in the format specification are allowed.
The format card is punched with parentheses surrounding the specifications in columns
1-80, as shown in Figure 3.
CC: 1
03, 11, F1.0, F5.2, F7.5, 3F2.1)
1
Figure 3. Format card example
3
Note: For each data card within an observation set (in case more than one
card is required per observation), there must be a variable format
card preceding the data deck. These format cards must be in the
same order as are the cards in the observation sets. At most, the
user will supply three format cards.
1.3 PROGRAM PAUSES AND MESSAGES
1. Pause 10. An illegal character in a numeric field has been encountered in reading
data. The program will print the card and the approximate column where the error
was detected. Pressing START on the card reader and console will cause the
remaining data cards to be read and ignored. The next monitor control card,
possibly signaling a new analysis, will be operated on. If this is not desired, the
following should be done:
It is possible that the format specification card is incorrect. If this is so, the entire
deck to be analyzed must be rerun. However, if a specific data card is in error, the
reader hopper and the stacker should be emptied. Pressing the nonprocess runout
button will clear the card read-punch, and the second card in the stacker will be the
card containing the error. After correcting this card, the user should place it and
the third stacker card at the front of the deck that was withdrawn from the hopper,
place this entire deck in the reader hopper, and press START on the card reader and
console to continue processing.
2. Other error conditions are signaled by a printed message, and/or the program exits
to the monitor. The monitor will read cards until a monitor control card is met
(that is, the next job to be done), or will stop when the reader hopper is empty. For
a list of the error messages, see chapter 5.
3. When an analysis is terminated successfully, an end-of-job message is printed, and
control is relinquished to the monitor.
4. When the user calls for output on cards, a message is written reminding the operator
to enter blank cards, if console entry switch 15 is not on. The computer then pauses
to allow input of blank cards. If console entry switch 15 is on, no reminder is given.
It is possible, in this case, to destroy the next analysis deck. See sections 1.4 and
1 .5.
1.4 STACKING: SEQUENTIAL PROGRAM OPERATION
Stacking of jobs is permitted. Each job must be a complete deck, as defined in the job
execution section of each program. However, when a program option card calls for
output on the IBM 1442 Card Punch, the negative identification card following the input
data must be succeeded by blank cards. For each matrix requested in factor analysis or
regression analysis, it is wise to place at least [n 2/5 + 2n + 2] blank cards behind the
data deck, where n is the number of variables processed. For orthogonal polynomials,
n + 1 blank cards should be included, where n is the order of the polynomial requested.
When factor scores are to be punched, 2n blank cards should be included, where n is
the number of observations processed.
4
It is advisable to place extra blank cards in the hopper, because an insufficient number
could result in the destruction of a part of the next analysis deck. After an analysis is
completed, cards are read until the next monitor control card is met.
1.5 TYPEWRITER AND PUNCHED CARD OUTPUT
For typewriter output, the 1130 Statistical System uses the same format statements as
are used for the printer. A user electing to use this device heavily may desire to
modify output using Assembler Language routines, calling on the typewriter tabulation
feature.
All system programs, excepting analysis of variance, allow user selection of punched
card output. This is discussed briefly in sections 1.3 (4) and 1.4. In the following,
a detailed explanation of the mechanics of this operation is given.
Consider first that stacking is not being done; only one job is being run. If the user
places one blank card behind his input deck, it is unnecessary to press the card reader
and console start buttons to complete the reading of the input data. If the user has asked
for punched output in his analysis definition (option card), an adequate number of blank
cards should be in the hopper following the data (section 1.4). If this is not the case, the
computer halts, waiting for the entry of blank cards. After these are placed in the
reader hopper, the start button should be pressed on the card reader and console to
continue processing.
If jobs have been stacked (section 1.4) and if, following the card with a negative
identification field signifying end of data, there is another analysis deck (the next job),
it is possible to destroy this next deck if punched output is being requested in the
current job.
In this case, if console switch 15 is down (off), a message is written reminding the
user to place blank cards in the hopper; then the computer pauses. If this occurs, the
card reader hopper (which contains the next job to be run) should be emptied, the non
process runout button on the card reader should be pressed, and the last two cards in
the stacker (//XEQ and a LOCAL card) should be placed at the front of the next job to
be run. Blank cards should then be placed in the reader hopper, followed by the next
job to be run, and START pressed on the card reader and console to continue processing.
1.6 MACHINE AND SYSTEMS CONFIGURATION
The 1130 Statistical System is designed to operate on an 8K 1130 Computing System with
disk storage (1131 Model II) and 1442 Card Read Punch; the 1132 Printer is optional.
It is written to operate under the 1130 Disk Monitor System (1130-0S-001).
1.7 PROGRAMMING LANGUAGE
IBM 1130 FORTRAN and the IBM 1130 Assembler Language.
5
1.8 REFERENCE MATERIAL
IBM 1130 Disk Monitor System Reference Manual (C26-3750)
IBM 1130 Assembler Language (C26-5927)
IBM 1130 FORTRAN Language (C26-5933)
6
CHAPTER 2: PROGRAMS
2.1 STEPWISE LINEAR REGRESSION
From sets of observations numbering 499 or fewer, containing measures on a
dependent variable y and n independent variables xl, x2, ...xn, where the total number
of variables is less than or equal to 30, the stepwise linear regression analysis will
determine the coefficients of a linear equation of the form:
y=b +bx +bx +...bx
22
nn
o
11
which best approximates the observations in the least-squares sense.
, xn are entered into the equation on the basis
The independent variables x l , x2 ,
of a variance criterion supplied by the user, which enables the program to determine
which variable makes the greatest improvement in "goodness of fit". Similarly.,
variables are not entered, or removed, from the equation on the basis of a second
variance criterion which indicates that the variable does not offer any significant
improvement in the goodness of fit.
The general method of solution to determine the coefficients b 0 , b 1 , ...bm is to
compute the matrix of correlation coefficients from the source data. This matrix will
contain the correlations between all the independent variables and the dependent
variable. By applying a Gaussian elimination inversion process, a stepwise inverse of
the correlation matrix is computed. Multiplying this inverse by a vector containing
the dependent variable correlated with each independent variable forms the normalized
regression coefficients. The inversion process is carried out for one variable at a time.
As each variable is processed, it is compared to the variance criterion to determine its
significance. If the variable is to be entered, the coefficients for the equation containing
a subset of the total number of variables in the analysis are computed and made
available for printout and use in the next step of the analysis. Because of the nature of
the computational process, the elements of several subsidiary statistics are also
available. If the user elects to print each regression step as it is computed, these
statistics will be printed with the regression coefficients.
The following book can be used as a reference: Ralston, A. and Wilf, H. S.
Mathematical Methods for Digital Computers. New York: John Wiley and Sons, Inc. ,
1960.
2.1.1 Summary of Output Statistics
1. High and low value of each variable
2. Means of each variable
3. Standard deviation of each variable
4. Sample variance for each variable
5. Matrix of raw cross products
7
6. Matrix of residual cross products
7. Variance-covariance matrix
8. Matrix of correlation coefficients
9. Residual Standard Deviation
10. Standard error of the mean of the predicted dependent variable
— 11.
Multiple correlation coefficient
— 12.
Square of the multiple correlation coefficient -
sum of squares due to regression
adjusted total sum of squares
13. Degrees of freedom
— 14. Regression coefficients - B
15. Standard error of regression coefficients
where n denotes the
rTinn
i
16. Partial correlation coefficients - r i = am /ai
dependent variable and a is an element of the stepwise inverse of the
correlation matrix.
; Bi = Bi Si / Sy where Si is the
17. Normalized regression coefficients standard deviation of the ith independent variable
18. Standard error of normalized regression coefficients
— 19. For each data case, the predicted value and difference between the predicted
value and the actual value
20. Analysis of variance table
2.1.2 Job Execution
To perform a regression analysis, the user must supply three sets of cards to the
program:
1. Monitor control cards
2. Program control cards
3. Data cards
Descriptions of the form and content of each card set follow.
MONITOR CONTROL CARDS
The monitor control cards are necessary to initiate program loading from the disk and
to establish the necessary communication with the monitor. A general description of
cards may be found in IBM 1130 Disk Monitor System Reference Manual (C26-3750).
8
A regression analysis requires the following cards:
16-17
CC: 1 4 8
// XEQ REGR 03
*LOCALREGR, FMTRD, PRNTB, DATRD, 1VDCRAD, TRAN
*LOCALCOREL, PRNT
*LOCALREGR2, REGRE
The monitor control cards do not change from job to job within one analysis type, but
must be included with every job processed. The first program operated on by this
system should be preceded by a cold start card.
PROGRAM CONTROL CARDS
The program control cards communicate the data-specific parameters and output
options to the program. There are five possible card types necessary for execution:
1. Input/output units card*
2. Job-title card*
3. Option card (described below)
4. Variable name card (described below)
5. Variable format card*
Four of the control cards are required in every job. The variable format card, which
specifies data format, is not necessary if matrix data is to be processed.
OPTION CARD
Number of Variables (cc 1-2)
This field must be punched with a nonzero integer, n, which is less than or equal to 30.
The value of integer n gives the total number of independent and dependent variables to
be processed.
Input Type and Source (cc 3-4)
This field allows the user to specify the input device (1442 card reader or disk) and,
indirectly, the type of input analysis to be undertaken in the input program. The three
possible values that may be punched in this field are described below:
Value
1
Meaning
Raw data will be read from the 1442 card reader and transferred
to the disk, where it will be retained until destroyed by input from
*See "General Operating Instructions", section 1.2.
9
Value
1
(cont)
Meaning
one of the four programs in this system. Until destroyed, option
2 (below) can be used to read this data from the disk.
2
Raw data will be read from the disk. Raw sums and raw sums of
cross products will be accumulated. Data will be read until a
negative number in the identification field is encountered
(section 2. 1.3).
3
A previously computed matrix, or matrices, will be read from
the 1442 card reader. Matrix cards will be read until a negative
job number field is encountered (see "Pooling", section 2.1.4).
Sequence Checking Within Observations (cc 5-6)
This field is used to indicate that raw data input from the card reader (cc 3-4 contains
a 1) is to be sequence-checked. A value of zero or a blank field implies that no
sequence check will be made. A value of one (1) implies that the cards will be
sequence-checked. The sequence-checking process consists of an equal comparison
check of the case identification field, for all cards in a case, and an ascending
sequence check of the card number field. If an error in either of these conditions is
encountered, the program prints a message, and the job is terminated.
Number of Variables on Card 1 (cc 7-8)
When a data vector contains more variables than will fit on one card, the user must
indicate to the program the number of variables punched on each card. This field must
be punched with the number of variables on the first card. If there is only one card
per case, this field must be blank or zero.
Number of Variables on Card 2 (cc 9-10)
Same as cc 7-8, except that this field indicates the number of variables on the second
card of the data.
Number of Variables on Card 3 (cc 11-12)
Same as cc 7-8, except that this field indicates the number of variables on the third
card of the data.
Transformation Switch (cc 13-14)
If the value in this field is nonzero, a user-written transformation subroutine is
called after each data record is read and before any computation takes place.
If the value in this field is zero or blank, the transformation subroutine is not called.
The use of transformations is discussed in section 2.5.1.
10
Output Raw Sums of Cross Products (cc 15-16)
This field is used to indicate whether the raw sums and sums of raw cross products
matrix are to be printed, punched, printed and punched, or not presented.
The four (4) possible values of this field are described below. The computation to
generate the matrix is performed even if the "no output" option is chosen.
Value
Meaning
0 or blank
No output.
1
Matrix will be printed.
2
Matrix will be printed and punched.
3
Matrix will be punched.
Punched output of the raw sums of cross products matrix includes the number of
observations and the vector of raw sums and sums of squares. This entire output
must be entered on the pooling option (section 2.1.3).
Output Residual Cross Products (cc 17-18)
This field is used to indicate whether the residual cross products matrix — defined as:
S .S.
= c..
1
n
i,j = 1,2...n
where c.. are the elements of sums of raw cross products matrix,
s. are the raw
, s.
1,
ij
th
th
sums of the i — and j — variables, respectively, and n is the number of cases — is to be
printed, punched, printed and punched, or not presented.
The four (4) possible values are described above under "Output Raw Sums of Cross
Products".
The matrix is computed even if the "no output" option is chosen.
Output Variance-Covariance Matrix (cc 19-20)
This field is used to indicate whether the variance-covariance matrix — defined as:
c.. =
1)
lJ
n-1
j--- 1,2.. . n
where uij is an element of the residual cross products matrix, and n is the number of
cases (or sum of weights) — is to be printed, printed and punched, punched, or not
presented. The four (4) possible values that may occur in this field are as given
for the above matrices.
11
There are no additional vectors or matrices punched with the punched output. The
matrix is computed even if the "no output" option is chosen.
Output Correlation Matrix (cc 21-22)
This field is used to indicate whether the correlation matrix — defined by:
i,j = 1,2...n
where c.. is an element of the variance-covariance matrix, and s., s. are the standard
ij
th
—th
deviations of the i and j — variables, respectively — is to be printed, punched, printed
and punched, or not presented. The four (4) possible values contained in this field
are given above under "Output Raw Sums of Cross Products".
The punched output of the correlation matrix includes the number of cases and cards
containing the vectors of means and standard deviations.
The matrix is generated even if the "no output" option is chosen.
Compute and Print Predicted Values (and Residuals) (cc 23-24)
This field is used to indicate the regression step to begin computing and printing the
predicted values, defined as:
Yi = b + b x + b x +
0
l il
2 i2
b x
k ik
th
where b b 1 , ..b are the coefficients of the regression equation for the k — step,
0' 1' •
k
and x.. are the source data elements.
ij
If this field is blank or zero, the predicted values are not computed. If the field is
negative, predicted values are printed only on the final step.
If it contains a nonzero value, k, the predicted values are computed for each regression
step equation containing k or more variables. For example, when k = 1, the predicted
values for all step equations are printed. When k = 3, the predicted values for the
equation containing three independent variables are printed. The printout also contains
the actual value of y and the difference between the predicted and the actual values.
If predicted values are requested when equations with fewer than p variables are not
desired, no predicted values are printed. That is, if column 23-24 contains a positive
integer less than the integer in column 25-26, no predicted values are printed.
12
Print Steps of Regression (cc 25-26)
If this field contains a value of k, all step equations containing k or more independent
variables are printed. For example, if all steps are desired, a value of one (1) forces
the printout of the equations containing 1,2, ...m independent variables.
If this field is zero or blank, no printout occurs. Only a correlation matrix is
calculated.
Pooling Option (cc 27-28)
When using the matrix input option (cc 3-4 are 03) and when pooling sums of squares
and cross products (section 2.1.3), if the user desires that matrices be subtracted
rather than added (of aid in deletion of outlyers), this field should be nonzero.
Number of the Dependent Variable (cc 29-30)
The regression analysis program uses the value punched in this field to rearrange the
correlation matrix, means, standard deviation, and variable names vectors, such that
the dependent variable always follows the independent variables. The user must
therefore indicate to the program the number of the dependent variable in this field.
The value punched must be greater than zero and less than or equal to 30. A value of
zero implies a regression analysis is not desired, and the program will exit after the
correlation analysis is complete.
Variance Criterion to Remove Variables (cc 31-34)
This field is used to determine whether an independent variable, when removed from
the equation, significantly increases the sample residual variance. The significance
of the increase is determined by comparing it to the variance criterion as punched in
this field. If the computed variance measure is greater than the criterion, the variable
is removed from the equation.
The form of the number to be punched is =Ex. A decimal point may replace any x.
If there is no decimal point, the number is taken to be . xxxx. The size of this
number depends on the information available to the particular analysis. If the
user has no idea about the size to be used, a number between .005 and .05 may
be acceptable.
It is possible for the user to set criterion levels for entry and removal of variables
that cause cycling of variables into and out of the model. The program does not check
this cycling possibility.
Variance Criterion to Enter Variables (cc 35-38)
This field has a similar function to the previous field, except that the value punched is
to determine whether a variable is to enter the equation. A variable is entered into
the equation if it significantly reduces the sample residual variance. The computed
variance is compared to this criterion value to determine whether it does decrease
this variance. The typical values and the form of the punched data are exactly as
described above for the remove-variable criterion.
13
A number twice the size of the removal factor can be tried if the user is not sure of the
correct number to be used in this field.
Tolerance For Ill Condition (cc 39-44)
A poorly conditioned matrix occurs when an independent variable is approximately a
linear combination of other independent variables. This number is associated, in the
program, with the size of the pivot element. If the pivot element is less than the
tolerance level, the associated variable is not entered into the model on the iteration
in which the condition occurs. A tolerance level of zero is not to be advised, unless
the user is sure that his matrix is not ill-conditioned.
Regression Analysis Option Card Summary
Column
Meaning
1-2
Number of variables
3-4
Input type and source
1 - Raw data input from card reader
2 - Raw data input from disk
3 - Matrix input from card reader
5-6
*Check sequence of raw data input
0 - No
1 - Yes
7-8
*Number of variables on card 1
9-10
*Number of variables on card 2
11-12
*Number of variables on card 3
13-14
*Transformation switch
0 - No transformation
1 - Transformation
(must be blank if there is only
one card per observation)
15-16
**Output raw cross products matrix
0 - No
1 - Print
2 - Print and punch
3 - Punch
17-18
**Output adjusted cross products matrix
0 - No
1 - Print
2 - Print and punch
3 - Punch
19-20
**Output variance-covariance matrix
0 - No
1 - Print
2 - Print and punch
3 - Punch
* Not pertinent when matrix input is used.
** When correlation matrix input is used, matrices are not available for output.
14
Column
21-22
23-24
Meaning
**Output correlation matrix
0 - No
1 - Print
2 - Print and punch
3 - Punch
*Output predicted values
1 - Print predicted values for last step only.
0 - Do not print predicted values.
k - Print predicted values for models containing k
or more independent variables
25-26
Output steps of regression
0 - Print no regression steps. Exit after correlation analysis.
k - Print all steps for models containing k or more independent variables.
27-28
Pooling option (see sections 2.5.3 and 2. 1.4)
Zero - Add matrices with ID = 1
Nonzero - Subtract matrices with ID = 1
29-30
Number of dependent variable
31-34
Variance criterion to remove variables
35-38
Variance criterion to enter variables
39-44
Tolerance for colinearity
VARIABLE NAME CARD (Figure 4)
In the multiple regression program there are a number of matrix printouts that the
user may request. The variables in the matrix may be assigned a four-character name
to aid in the identification of the output. The card is punched in four-column fields, and
each field corresponds to the variable to be identified (for example, field 3 (columns
9-12) will be the name of row and column 3 on all matrix output). At most, 20 names
can appear on one card. If there are more than 20 variables in the analysis, a second
card having the same format as the first must be included in the control card deck.
Column
Meaning
1-4
Name of variable 1.
5-8
Name of variable 2.
(4N-3) - (4N) Name of variable N.
* Not pertinent when matrix input is used.
** When correlation matrix input is used, matrices are not available for output.
15
CC: 1
4 5
8 9
'GRP 1 ANL
12 13 16
XX
Figure 4. Variable name card example
2.1. 3 Data Input
Raw data input to the regression program consists of a set of observations made on
several different variables. The variables for each observation are punched on one,
two, or three cards, according to the following general form:
Field
Type
Meaning
1
Integer (I)
Identification field. Any numeric information that serves
to identify the particular observation is punched in this
field. It must be greater than zero, and should be
different for each observation.
2
Integer (I)
Card number within observation. If it is not possible
for one card to contain all the variables, they may be
continued on a second and a third card, as necessary.
The user has the option of sequence checking the cards
to ensure that all cards within a case are together, and
that the order of cards is consistent. If the option is
chosen (cc 5-6 on option card), this field must be punched
with an integer that is in ascending sequence for all cards
in the case. If sequence checking is not desired, the
field may be blank and may consist of one blank column.
3, 4, . . . , n Floating
etc.
point (F)
Observation on variable x l . Any number may be
punched in this field. Decimal points are not required.
The remaining fields on the card are reserved for variable
observations. If there are more variables than can fit on the
first card, a second and a third card may be used.
The particular card columns for each field are arbitrary.
16
Following the data deck, the user must include a card containing a negative integer in the
identification field. This card signals the end of data.
2.1.4 Matrix Input/Output
It is possible to obtain punched card output of a number of matrices (see section 2.5.3)
and vectors with the regression program. This program is designed to also input some
of these matrices, at a later time, for further analysis or processing. In addition,
matrices from another program or source, if punched in the proper format, may
also be used as input.
This section is devoted to a description of various possible forms of analysis with the
output options available in each program.
Format Description
Matrices are punched rowwise, five elements to a card, in the FORTRAN E or floatingpoint format. Each card is identified as to its job number, matrix number, row
number, and column number of the first element on the card. The specific card
columns occupied by the identification fields, and matrix elements, are shown below:
Column
Meaning
1-4
Job number
5-6
Matrix identification
number (section 2.5.3)
7-8
Column number of first
element on card
9-10
Row number
11-24
Matrix element
(±0. XXXXXXXE±NN)
25-38
Matrix element
39-52
Matrix element
53-66
Matrix element
67-80
Matrix element
Note that all matrices are punched and read under a fixed format. Hence, a variable
format card is not allowed when using punched card matrices as input.
Most matrices have a unique identification number. However, there are a few cases
where two or three vectors have the same identification and are always punched
together. For these cases, see section 2.5.3.
17
Regression with Correlation Matrix Input
The punched output option of the correlation matrix includes the punchout of the number
of cases (matrix 21), and means and standard deviation vectors (matrix 23). This
complete output can be used as input to initiate another analysis without the necessity
of reprocessing the raw data used to generate the matrices.
To use the correlation matrix set as input, the user places the punched output behind
the variable names card, followed by a card that contains a negative number in the job
number field. The program reads the number of cases, means, standard deviations,
and correlation matrix, storing each in its appropriate location, and, then initiates the
analysis as specified on the option card. No matrix output is possible in this case.
Also, observations are not read; hence, predicted values and certain summary statistics
are not available.
Pooling Sums of Squares and Cross Products (cc 27-28)
In the regression and factor analysis programs, a considerable amount of processing
time is devo,ted to accumulating raw sums and raw cross products as each data vector
is read. If there is a large amount of source data, or if there is some logical
division in the data set, it is frequently desirable to obtain partial punched output of the
raw sums of squares and cross products. These partial outputs can then be added
together to complete the total analysis in another job.
Both programs allow this type of analysis. By choosing the punchout option for this
matrix, the program includes in the punchout the number of cases, and raw sums
and sums of squares vectors, in addition to the raw cross products matrix. Any
number of these matrices may be punched and used later to complete the analysis. The
user simply stacks each output set, one after the other, following the variable names
card. The program reads the matrices, examining the matrix identification and row
and column numbers to determine the location or group of locations to which the matrix
is to be added. The read-add operation is terminated when a card with a blank or
negative job number field is encountered, unless the pooling option (cc 27-28 of the
option card) is nonzero. In this case, the read-add operation terminates at the first
detection of a blank or negative job number field, and the second (succeeding) matrix
is subtracted from the previous matrix. This operation is terminated when the second
blank or negative job number is encountered.
If the subtraction option is used, the second set of matrices must also include its
associated raw sums and sums of squares vectors for proper analysis completion.
Predicted values are not available with this option. Also, high and low values are not
calculated for the observations on variables.
If outlyers are detected, the user has two options available if he wishes to reanalyze
ignoring these outlyers:
1. He can eliminate the data cards containing the outlyers, and rerun the entire
analysis.
18
2. He can prepare cards according to the format given above under "Format
Description", either by hand or by using the program. To use the program, he must
run the analysis using only data cards associated with outlyers. The option card must
request raw cross products matrix output, and may note that the dependent variable is
zero, so that the analysis will terminate after the correlation matrix is calculated.
If the user allows an entire regression analysis to be computed using only the outlyer
cards, a termination with some error condition may result — for example, mean
square nonpositive. In any case, matrices 1, 21, and 22 will be punched (see
section 2.5.3). These matrices should be used as the second set of matrices for
input using the subtraction option.
2.1.5 Operating Instructions
A.
Using the regression analysis program when the total 1130 Statistical System has
not been stored on the disk
If the user wishes to load only the set of programs that allow regression analyses, the
following programs must be compiled or assembled and stored on the disk. Each deck
begins with a card punched as
//FOR
and ends with an
*STORE
card.
The user should use a disk containing the 1130 Disk Monitor System, as described in
section 1.1. The following decks should be preceded by a cold start card, placed in
the card reader hopper, and the buttons IMMEDIATE STOP (console), RESET (console),
START (card reader), and PROGRAM LOAD (console) should be pressed. A blank card
should be placed after the last deck in the card reader hopper.
DECKS-LABELS: REGR-REGR; **COREL-CORL; **PRNT-PRNT;
*FMTRD-FMRD; *DATRD-DTRD; *PRNTB-PRNB; *GMPYX-GMPY;
*GDIVX-GDIV; **IVEXRAD-MXRD; REGR2-RGR2; REGRE-RGRE;
*FMAT-FMAT; TRAN-TRAN.
In addition, regression and factor analysis programs must reside on the disk together;
section 2.2.7 names additional routines to be placed on the disk.
B.
Execution from disk
Once the component subroutines and main calling programs are on the disk, the execution
of a job requires the monitor control cards, program control cards, and data cards to be
*Used in all four analysis types
**Used in factor analysis
19
placed in the card reader. The deck should be preceded by a cold start card. To
initiate processing, the buttons IMMEDIATE STOP and RESET (console), START (card
reader and printer), and PROGRAM LOAD (console) should be pressed. The order in
which the cards are placed in the card reader for either matrix or raw data input is
shown in Figures 5, 6, and 7.
ariable Forma
(Variable Names
(Option Card
(Job-Title Card
Input/Output
( Units
Monitor
Control Cards
Figure 5. Regression card order — card reader input
Optional
Blank Output
/Variable
Names
(Option
(Job-Title
/ Input/Output
Units
Monitor
Control
Figure 6. Regression card order — disk input
20
Optional Blank
Output
Negative
Identification
((
Matrix
..e(to be added to matrix 1)
Variable
( Names
(Option
(Job-Title
(I/O Units
Monitor
Control
Figure 7. Regression card order — matrix input
2. 1. 6 Sample Problem
INPUT
// XEQ REGR
03
*LOCALREGR,FMTRO/PRNT8.DATRO,MXRAD,TRAN
*LOCALREGR2,REGRE
*LOCALCORELPRNT
020200
STEPWISE TEST ONE
2222
0001010102-1010006.500.300.00010
06010000
P1 P2 P3 P4 P5 P6
(212 9 1X .F5.2,F5.0,2F5.2,2F5.0)
0101 002500002502500001500003400064
0201 013000002102100000870003600065
0301 003500002202200000430004100082
0401 001750000900130001800001500023
0501 003000002302300002000003300064
0601 002000001000060003300001300016
0701 005500000700140003400001600012
0801 006000000600080005000001100027
0901 001300000800270001500001900048
10.01 005000001800360001800002700050
1101 005000000300100001400001400012
1201 003000000800270001000002500013
1301 002000000600300001500002100020
1401 002000000800100002500001800023
1501 001000002202200001100004600118
1601 004000001301300002800001700050
1701 000500002600120000730004800063
1801 000250002302300000100003600150
1901 014000000300100003500000500072
2001 002500001500250000280003300054
2101 003500002801400000010004600109
2201 003500000600060005000001000010
2301 002500003503500005700003800125
2401 000500001100200003400001600044
2501 002000001101100000500002000048
2601 007000003203200006600003800105
2701 004000000800100004500001200009
2801 015000002302300000150004900130
2901 001000003803800002200004300160
3001 003500001500500001500003300048
21
3101
3201
3301
3401
3501
3601
3701
3801
3901
4001
4101
4201
4301
4401
4501
4601
4701
4801
4901
5001
5101
5201
5301
5401
5501
5601
5701
5801
5901
6001
6101
6201
6301
6401
6501
6601
6701
6801
-1
013000000600120003700000900036
002000002502500001000003500150
012000000500170000300002100078
004000000900075001900001700023
003000000700350002600001200042
008000002002000002200003000072
009000000600086002500001500020
006000001200400001200002000036
008000002600160001100003500056
001500001500300001600002900036
007000001000090010000001200026
008000002802800004200004000108
002000003403400000900004200106
006000000400080003600001100016
015000003203200001800004400104
017000001101100002300001400047
016000000200050001800001100027
003000001800160001100003200012
006000000300040001300001500007
014000000800110002000001700018
006000001400090000700002900028
001800001200240001500002100025
015000000300150000800001300011
018000000600550005700000900020
005000001200200004100001600014
030000001101100002000002200038
029000000800800001000002200103
001800002402400001100003800106
013000002602600001700003800063
019000002902900048000002900208
011000001701700001600002500032
010000001500500003500001900028
006000001000500001000002600032
005000002202200001200003900100
001000001500500000800002900050
017000000900300013000001000080
005000003003500000900005800065
001300001000130009000001000025
PUNCHED CORRELATION MATRIX OUTPUT
2222
2222
2222
2222
2222
2222
4 1
4 1
4 1
4 1
4 1
4 1
2222 4 6
2222 4 6
2222 4 6
2222 4 6
2222 4 6
2222 4 6
222223 1
222223 1
222223 1
222223 1
222223 1
222223 1
222221 1
1 0.1000000E C1-0.1764650E
2-0.1764650E 00 0.1000000E
3 0.5134508E-02 0.8679906E
4 0.2554817E 00 0.1007193E
5-0.1956896E 00 0.8791197E
6 0.8205512E-01 0.7496167E
1 0.8205512E-01
2 0.7496167E 00
3 0.7860574E 00
4 0.3425673E 00
5 0.6451673E 00
6 0.1000000E 01
1 0.6995588E 01 0.6473750E
2 0.1525000E 02 0.9357534E
3 0.1042514E 02 0.1162702E
4 0.3099554E 01 0.5997127E
5 0.2539706E 02 0.1247940E
6 0.5679412E 02 0.4355013E
1 0.6800001E 02
OUTPUT
22
00
01
00
00
00
00
01
01
02
01
02
02
0.5134508E-02 0.2554817E
0.8679906E 00 0.1007193E
0.100000OE 01 0.1258658E
0.1258658E 00 0.1000000F
0.7519358E 00-0.1404377E
0.7860574E 00 0.3425673E
00-0.1956896E
CO 0.8791197E
00 0.7519158E
01-0.1404 17 7E
00 0.100000OF
00 0.6451673E
CD
CD
v.
CD
01
CO
// XEO REGR
03
•LOCALREGR.FMTRDORNT0.DATRDoMXRADORAN
•LOCALREGR2.REGRE
•LOCALCORELORNT
JOB 2222
STEPWISE TEST ONE
0
PAGE
6
NUMBER OF VARIABLES
1
INPUT TYPE
0
SEQUENCE CHECK
0
VARIABLES ON CARD 1
0
VARIABLES ON CARD 2
0
VARIABLES ON CARD 3
0
TRANSFORMATION SWITCH
1
OUTPUT RAW CROSS PRODUCTS
1
OUTPUT RESIDUAL CROSS PRODUCTS
PRINT PREDICTED VALUES
-1
1
PRINT STEPS
0
POOLING OPTION
6
DEPENDENT VARIABLE
0.500
F-LEVEL TO REMOVE VARIABLES
0.300
F-LEVEL TO ENTER VARIABLES
0.00010
TOLERANCE VALUE
1
OUTPUT VARIANCE ■ COVARIANCE
2
OUTPUT CORRELATION
(212.1X .F5.2.F5.0.2F5.2.2F5.01
STEPWISE TEST ONE
JOB 2222
PAGE
1
PAGE
2
MATRIX OF RAW CROSS-PRODUCTS
VARIABLE
01
P2
P3
P4
PS
P6
P1
0.61357E
0.65381E
0.49851E
0.21390E
0.11022E
0.28566E
04
04
04
04
05
05
P2
0.65381E
0.21681E
0.17138E
0.35929E
0.33215E
0.79363E
04
05
05
04
05
05
P3
0.49851E
0.17138E
0.16448E
0.27853E
0.25314E
0.66929E
04
05
05
04
05
05
P4
0.21390E
0.35929E
0.27853E
0.30629E
0.46487E
0.17964E
04
04
04
04
04
05
P5
0.11022E
0.33215E
0.25314E
0.46487E
0.54295E
0.12157E
05
05
05
04
05
06
STEPWISE TEST ONE
P6
0.28566E
0.79363E
0.66929E
0.17964E
0.12157E
0.34641E
05
05
05
05
06
06
JOB
2222
MATRIX OF RESIDUAL CROSS-PRODUCTS
VARIABLE
P1
P2
P3
P4
PS
P6
P1
0.28079E
■0.71622E
0.25893E
0.66455E
■0.10592E
0.15499E
P2
04 ■0.71622E 03
03 0.58667E 04
02 0.63273E 04
03 0.37869E 03
04 0.68782E 04
04 0.20467E 05
P3
0.25893E
0.63273E
0.90575E
0.58802E
0.73100E
0.26667E
P4
P5
02 0.66455E 03
.10592E 04
04 0.37869E 03 0.68782E 04
04 0.5ele02E 03 0.73100E 04
03 0.24096E 04 ■0.70419E 03
04 ■0.70419E 03 0.10434E 05
05 0.59945E 04 0.23492E 05
P6
0.15499E
0.20467E
0.26667E
0.59945E
0.23492E
0.12707E
04
05
05
04
05
06
23
STEPWISE TEST ONE
JOB
2222
PAGE
3
VARIANCE - COVARIANCE MATRIX
VARIABLE
P1
P2
P3
P4
P5
P6
P1
0.41909E
*0.10689E
0.38647E
0.99187E
-0.15809E
0.23134E
P2
02 *0.10689E 02
02 0.87563E 02
00 0.94437E 02
01
0.56521E 01
02 0.10266E 03
02 0.30548E 03
P3
0.38647E
0.94437E
0.13518F
0.87764E
0.10910E
0.39802E
P4
P5
00 0.99187E 01 *0.15809E 02
02 0.56521E 01 0.10266E 03
03
0.87764E 01 0.10910F 03
01
0.35965E 02 .-0.10510E 02
03 *0.10510E 02 0.15573E 03
03 0.89470E 02 0.35063E 03
P6
0.23134E
0.30548E
0.39802E
0.8947CE
0.35063E
0.18966E
STEPWISE TEST ONE
JOB
SUMMARY STATISTICS
P1
P2
P3
P4
P5
P6
PAGE
2222
NO.OF CASES •68
VAR/ABLE
1
2
3
4
5
6
02
03
03
02
03
04
LOW
0.25000E 00
0.20000E 01
0.40000E 00
0.10000E-01
0.50000E 01
0.70000E 01
HIGH
0.30000E
0.38000E
oos000E
0.48000E
0.58000E
0.20800E
AVERAGE
02
02
02
02
02
03
0.69955E
0.15250E
0.10425E
0.30995E
0.25397E
0.56794E
STD. DEV.
01
02
02
01
02
02
0.64737E
0.93575E
0.11627E
0.59971E
0.12479E
0.43550E
01
01
02
01
02
02
VARIANCE
0.41909E
0.87563E
0.13518E
0.35965E
0.15573E
0.18966E
02
02
03
02
03
04
READY THE PUNCH WITH BLANK CARDS AND PRESS START ON THE PUNCH AND CONSOLE. TURN CONSOLE SWITCH 15 ON.
STEPWISE TEST ONE
JOB 2222
MATRIX OF CORRELATION COEFFICIENTS
VARIABLE
P1
P2
P3
P4
P5
P6
24
PI
P2
0.10000E 01 *0.17646E 00
-0.17646E 00 0.10000E 01
0.51345 -02 0.86799E 00
0.25548E 00 0.10071E 00
0.19568E 00 0.87911E 00
0.82055E*01
0.74961E 00
P4
P3
P5
0.51345 -02 0.25548E 00 -0.19568E 00
0.86799E 00 0.10071E 00 0.87911E 00
0.10000E 01 0.12586E 00 0.75193E 00
0.12586E 00 0.10000E 01 *0.14043E 00
0.75193E 00 *0.14043E 00 0.10000E 01
0.78605E 00 0.34256E 00 0.64516E 00
P6
0.82055E*01
0.74961E 00
0.78605E 00
0.34256E 00
0.64516E 00
0.10000E 01
PAGE
4
STEPWISE TEST ONc
JOB
2222
PAGF
6
REGRESSION ANALYSIS
DEPENDENT VARIABLE
RESIDUAL STANDARD DEVIATION
STANDARD ERROR OF THE MEAN
MULTIPLE R
MULTIPLE RSOR
P6
27.1238
3.2892
0.7860
0.6178
VARIABLE ENTERED
VARIABLE
P3
B
P3
.OEF
STD ERROR OF B
2.9442
CONSTANT
PARTIAL-R
0.2850
0.7860
BETA-COEF
STD ERROR OF BETA
0.7860
0.0760
26.0998
ANALYSIS OF VARIANCE TABLE
SOURCE
D.E.
MEAN
REGRESSION
ERROR
1
1
66
SUM OF SQUARES
0.21933E U6
0.78516E 05
0.48556E 05
MEAN SQUARE
0.21933E 06
0.78516E 05
0.73570E 03
0.10672E 03
STEPWISE TEST ONE
REGRESSION ANALYSIS
P6
25.0821
3.0416
0.8235
0.6781
DEPENDENT VARIABLE
RESIDUAL STANDARD DEVIATION
STANDARD ERROR OF THE MEAN
MULTIPLE R
MULTIPLE RSOR
P4
VARIABLE ENTERED
VARIABLE
'OEF
STD ERROR OF B
2.8275
1.7976
P3
P4
0.2656
0.5150
PARTIAL-R
0.7971
0.3972
BETA - COEF
0.7548
0.2475
STD ERROR OF BETA
0.0709
0.0709
21.7445
CONSTANT
ANALYSIS OF VARIANCE TABLE
SOURCE
D.P.
MEAN
REGRESSION
ERROR
2
65
1
SUM OF SQUARES MEAN SQUARE
0.21933E 06
0.86180E OS
0.40892E 05
0.21933E 06
0.43090E OS
0462911E 03
F
0.68493E 02
25
STEPWISE TEST ONE
J08 2222
PAGE
8
REGRESSION ANALYSIS
DEPENDENT VARIABLE
RESIDUAL STANDARD DEVIATION
STANDARD ERROR OF THE MEAN
MULTIPLE R
MULTIPLE RSQR
P6
23.9328
2.9022
0.8435
0.7115
VARIABLE ENTERED
P5
VARIABLE
B ■ COEF
P3
P4
P5
STD ERROR
1.9583
2.3123
1.0355
CONSTANT
OF
B
0.4079
0.5266
0.3808
PARTIAL■R
BETA-COEF
0.5145
0.4811
0.3217
0.5228
0.3184
0.2967
STD ERROR OF BETA
0.1089
0.0725
0.1091
2.9105
ANALYSIS OF VARIANCE TABLE
SOURCE
D.F.
MEAN
REGRESSION
ERROR
3
64
SUM OF SQUARES MEAN SQUARE
0.21933E 06
0.90415E 05
0.36658E 05
1
0.21933E 06
0.30138E 05
0.57278E 03
0.52617E 02
STEPWISE TEST ONE
JOB 2222
PAGE
REGRESSION ANALYSIS
DEPENDENT VARIABLE
RESIDUAL STANDARD DEVIATION
STANDARD ERROR OF THE MEAN
MULTIPLE R
MULTIPLE RSOR
P6
23.9726
2.9071
0.8456
0.7150
VARIABLE ENTERED
VARIABLE
P1
B ■ COEF
P1
P3
P4
P5
STD ERROR OF
0.4272
1.8967
2.2333
1.1167
CONSTANT
0.4813
0.4145
0.5349
0.3923
PARTIAL■R
BETA-COEF
0.1111
0.4994
0.4654
0.3375
0.0635
0.5063
0.3075
0.3200
-1.2530
ANALYSIS OF VARIANCE TABLE
26
SOURCE
D.F.
MEAN
REGRESSION
ERROR
1
4
63
SUM OF SQUARES MEAN SQUARE
0.21933E 06
0.90867E 05
0.36205E 05
0.21933E 06
0.22716E 05
0.57468E 03
F
0.39529E 02
STD ERROR OF BETA
0.0715
0.1106
0.0736
0.1124
9
STEPWISE TEST ONE
JOB
2222
PAGE
10
PREDICTED VALUES
ACTUAL
CASE
0.6400E
0.6500E
0.8200E
0.2300E
0.6400E
0.1600E
0.1200E
0.27005
0.48005
0.5000E
0.1200E
0.13005
0.2000E
0.2300E
0.1180E
0.5000E
0.6300E
0.1500E
0.7200E
0.5400E
0.1090E
0.1000E
0.1250E
0.4400E
0.4800E
0.1050E
0.9000E
0.1300E
0.1600E
0.4800E
0.3600E
0.1500E
0.7800E
0.2300E
0.4200E
0.7200E
0.2000E
0.3600E
0.5600E
0.3600E
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
PREDICTED
02
02
02
02
02
02
02
02
02
02
02
02
02
02
03
02
02
03
02
02
03
02
03
02
02
03
01
03
03
02
02
03
02
02
02
02
02
02
02
02
mosesE
0.6627E
0.8871E
0.2275E
0.8497E
0.2262E
0.2921E
0.2627E
0.2899E
0.4188E
0.2154E
0.3530E
0.3209E
0.2718E
0.9473E
0.5035E
0.56475
0.8290E
0.2002E
0.4203E
0.7818E
0.2371E
0.1213E
0.2821E
0.4391E
0.1196E
0.2580e
0.1038E
0.1241E
0.4992E
0.2489E
0.8833E
0.3121E
0.25105
0.25875
0.7851E
0.2655E
0.3391E
0.46745
0.4103E
RESIDUAL
02
02
02
02
02
02
02
02
02
02
02
02
02
02
02
02
02
02
02
02
02
02
03
02
02
03
02
03
03
02
02
02
02
02
02
02
02
02
02
02
-0.2455E 02
.0.2127E 02
.0.6717E 01
0.2682E 00
.0.2097E 02
■.0.6627E 01
■0.1721E 02
0.72135 00
0.1900E 02
0.8116E 01
-0.9541E 01
■0.2230E 02
0.1209E 02
■0.4183E 01
0.2326E 02
.0.3517E 00
0.65285 01
0.6709E 02
0.5197E 02
0.1196E 02
0.3081E 02
■0.1371E 02
0.3631E 01
0.1578E 02
0.4082E 01
■0.1461E 02
.0.1680E 02
0.2616E 02
0.3581E 02
0.19243E 01
0.11105 02
0.61665 02
0.4678E 02
■0.2106E 01
0.1612E 02
■0.6515E 01
.0.6557E 01
0.20875 01
0.92575 01
0oo5037E 01
STEPWISE TEST ONE
JOB
2222
PAGE
11
PREDICTED VALUES
CASE
41
42
43
44
45
46
47
48
49
SO
51
52
S3
54
SS
56
57
58
59
60
61
62
63
64
65
66
67
66:
ACTUAL
0.2600E
0.1080E
0.1060E
0.1600E
0.1040E
0.4700E
0.2700E
0.1200E
0.7000E
0.1800E
0.2800E
0.2500E
0.1100E
0.2000E
0.1400E
0.3600E
0.1030E
0.1060E
0.6300E
0.2080E
0.3200E
0.2800E
0.3200E
0.1000E
0.5000E
0.8000E
0.6500E
0.25005
PREDICTED
02
03
03
02
03
02
02
02
01
02
02
02
02
02
02
02
03
03
02
03
02
02
02
03
02
02
02
02
0.3917E
0.1093E
0.1130E
0.2315E
0.1190E
0.4764E
0.2283E
0.41255
0.2172E
0.3026E
0.3696E
0.3087E
0.24305
0.3964E
0.3170E
0.6146E
0.5311E
0.8993E
0.99845
0.2014E
0.6718E
0.41535
0.4206E
0.8884E
0.4283E
0.5190E
0.1340E
0.3303E
02
03
03
02
03
02
02
02
02
02
02
02
02
02
02
02
02
02
02
03
02
02
02
02
02
02
03
02
RESIDUAL
0.1317E
■0.1323E
.0.7004E
■0.7151E
.0.1500E
■0.6450E
0.4164E
0.2925E
.43.1472E
■0.1226E
-0.8966E
.0.58695
■0.1330E
■0.1964E
.0.1770E
■0.2346E
0.4988E
0.1606E
.0.36134E
0.65435
-0.3518E
.0.1353E
..0.1006E
0.1115E
0.7169E
0.2809E
'.0.6905E
■0.8035E
02
01
01
01
02
00
01
02
02
02
01
01
02
02
02
02
02
02
02
01
02
02
02
02
01
02
02
01
JOB COMPLETED
27
CORRELATION MATRIX INPUT
// XEQ REGR
03
*LOCALREGR,FMTRDORNTBOATRDIMXRAD,TRAN
*LOCALREGR2,REGRE
*LOCALCOREL,PRNT
020200
STEPWISE TEST ONE
2222
06030000
000000000000010006.500.300.00010
P1 P2 P3 P4 P5 P6
2222 4 1 1 0.1000000E 61-0.1764650E 00 0.5134508E-02 0.2554817E
2222 4 1 2-0.1764650E 00 0.1000000E 01 0.8679906E 00 0.1007193E
2222 4 1 3 0.5134508E-02 0.8679906E 00 0.1000000E 01 0.1258658E
2222 4 1 4 0.2554817E 00 0.1007193E 00 0.1258658E 00 0.1000000E
2222 4 1 5-0.1956896E 00 0.8791197E 00 0.7519358E 00-0.1404377E
2222 4 1 6 0.8205512E-01 0.749616TE 00 0.7860574E 00 0.3425673E
2222 4 6 1 0.8205512E-01
2222 4 6 2 0.7496167E 00
2222 4 6 3 0.7860574E 00
2222 4 6 4 0.3425673E 00
2222 4 6 5 0.6451673E 00
2222 4 6 6 0.1000000E 01
222223 1 1 0.6995588E 01 0.6473750E 01
222223 1 2 0.1525000E 02 0.9357534E 01
222223 1 3 0.1042514E 02 0.1162702E 02
222223 1 4 0.3099554E 01 0.5997127E 01
222223 1 5 0.2539706E 02 0.1247940E 02
222223 1 6 0.5679412E 02 0.4355013E 02
222221 1 1 0.6800001E 02
-1
00-0.1956896E
00 0.8791197E
00 0.7519358E
01-0.1404377F
00 0.1000000E
0.0 0.6451673E
03
CO
00
03
01
00
OUTPUT
03
// XEO REGR
• LOCALREGRoFMTRD.PRNTBOATRD.MXRAD•TRAN
•LOCALREGR2oREGRE
•LOCALCORELORNT
STEPWISE TEST ONE
6
NUMBER OF VARIABLES
3
INPUT TYPE
0
SEQUENCE CHECK
0
VARIABLES ON CARD 1
0
VARIABLES ON CARD 2
0
VARIABLES ON CARD 3
0
TRANSFORMATION SWITCH
0
OUTPUT RAW CROSS PRODUCTS
0
OUTPUT RESIDUAL CROSS PRODUCTS
0
PRINT PREDICTED VALUES
1
PRINT STEPS
0
POOLING OPTION
6
DEPENDENT VARIABLE
0.500
F.-LEVEL TO REMOVE VARIABLES
0.300
F-LEVEL TO ENTER VARIABLES
0.00010
TOLERANCE VALUE
0
OUTPUT VARIANCE - COVARIANCE
0
OUTPUT CORRELATION
28
JOB 2222
PAGE
0
JOB
STEPWISE TEST ONE
2222
PAGE
1
REGRESSION ANALYSIS
OFPFNDENT VARIABLE
RFS1DUAL STANDARD DEVIATION
STANDARD ERROR OF THE MFAN
MULTIPLE R
MULTIPLE RUM
P6
27.1238
3.2892
0.7860
0.6178
VARIABLE ENTERED
VARIABLE
COEF
P3
STD ERROR OF B
2.9442
0.2850
PARTIAL-R
0.7860
BETA-COEF
0.7860
STD ERROR OF BETA
0.0760
26.0998
CONSTANT
ANALYSIS OF VARIANCE TABLE
SOURCE
D.P.
MEAN
4EGRESSION
ERROR
66
SUM OF SQUARES MEAN SQUARE
0.21933E 06
0.78516E 05
0.48556E 05
1
1
0.21933E 06
0.78516E 05
0.73570E 03
0.10672E 03
JOB 2222
STEPWISE TEST ONE
PAGE
2
REGRESSION ANALYSIS
P6
25.0821
3.0416
0.8235
0.6781
DEPENDENT VARIABLE
RESIDUAL STANDARD DEVIATION
STANDARD ERROR OF THE MEAN
MULTIPLE R
MULTIPLE (MR
P4
VARIABLE ENTERED
B ■ COEF
VARIABLE
P3
2.8275
P4
1.7976
CONSTANT
STD ERROR OF B
0.2656
0.5150
PARTIAL■R
BETA-COEF
0.7971
0.3972
0.7548
0.2475
STD ERROR OF BETA
0.0709
0.0709
21.7445
ANALYSIS OF VARIANCE TABLE
SOURCE
D.F.
MEAN
REGRESSION
ERROR
1
2
65
SUM OF SQUARES MEAN SQUARE
0.21933E 06
0.86180E 05
0.40892E 05
0.21933E 06
0.43090E 05
0.62911E 03
0.68493E 02
29
JOB
STEPWISE TEST ONE.
2222
PAGE
3
REGRESSION ANALYSIS
DEPENDENT VARIABLE
RESIDUAL STANDARD DEVIATION
STANDARD ERROR OF TMF. MEAN
M ULTIPLE R
MULTIPLE RSOR
P6
23.9328
2.9022
0.8435
0.7115
VARIABLE ENTERED
VARIABLE
P5
B ■ COEF
P3
P4
P5
STD ERROR OF B
0.4079
0.5266
0.3808
1.9583
2.3123
1.0355
CONSTANT
PARTIAL-R
0.5145
0.4811
0.3217
BETA-COEF
STD ERROR OF BETA
0.5228
0.3184
0.2967
0.1089
0.0725
0.1091
2.9105
ANALYSIS OF VARIANCE TABLC
SOURCE
D.F.
MEAN
REGRESSION
ERROR
3
64
SUM OF SQUARES MEAN SQUARE
0.21933E 06
0.90415E 05
0.36658E 05
1
0.21933E 06
0.30138E 05
0.57278E 03
0.52617E 02
STEPWISE TEST ONE
JOB
2222
PAGE
REGRESSION ANALYSIS
DEPENDENT VARIABLE
RESIDUAL STANDARD DEVIATION
STANDARD ERROR OF THE MEAN
MULTIPLE R
MULTIPLE RSOR
P6
23.9726
2.9071
0.8456
0.7150
VARIABLE ENTERED
VARIABLE
P1
B ■ COEF
PI
P3
P4
P5
STD ERROR OF B
0.4272
1.8967
2.2333
1.1167
CONSTANT
0.4813
0.4145
0.5349
0.3923
PARTIAL■R
BETA-COEF
0.1111
0.4994
0.4654
0.3375
0.0635
0.5063
0.3075
0.3200
-1.2530
ANALYSIS OF VARIANCE TABLE
JOB COMPLETED
30
SOURCE
D.F.
MEAN
REGRESSION
ERROR
1
4
63
SUM OF SQUARES MEAN SQUARE
0.21933E 06
0.90867E 05
0.36205E 05
0.21933E 06
0.22716E 05
0.57468E 03
0.39529E 02
STD ERROR OF BETA
0.0715
0.1106
0.0736
0.1124
4
2.2 PRINCIPAL COMPONENTS AND FACTOR ANALYSIS
The aim of factor analysis is to explain observed relationships among numerous
variables in terms of simpler relations. This simplification can take the form of
producing a set of classificatory categories, or creating a smaller set of hypothetical
variables.
The usual procedure is to collect measurements on n variables, over N persons or
objects (N should be appreciably larger than n). To find out "what goes with what"
among these n variables, the n variables can be intercorrelated as they vary over the
N objects. This is done for all possible n(n-1)/2 pairings of the variables, producing
a square symmetrical correlation matrix, R.
The process of factor analysis (or principal component analysis) is designed to resolve
this correlation matrix into an n x k factor matrix, in which the number of factors, k,
is usually considerably smaller than n, the number of variables. These factors may be
considered as underlying influences which, in further measurement, can be substituted
for the more numerous original variables, and which largely account for the correlations among the latter.
In analyzing the structure of a correlation matrix, two approaches can be taken.
Formally, they resemble one another to a certain extent, but they have, in fact, rather
different aims. One method is principal component analysis; the other is factor analysis.
The former method is a relatively simple technique of "breaking down" a correlation
matrix into a set of orthogonal (uncorrelated) components equal in number to the original
variables. These correspond to the latent roots (eigenvalues) and accompanying latent
vectors (eigenvectors) of the matrix. The method has the property that the roots are
extracted in descending order of magnitude; this is important if only a few of the
components are to be used for summarizing the data. These vectors are mutually
orthogonal, and the components derived from them are uncorrelated. Although a few
components may extract a large proportion of the total variance of the original variables,
all components are required to reproduce the correlations between the variables
exactly. Note that when the principal components method is employed, no hypothesis
need be made about the original variables. They need not even be random variables.
although, in practice, their values are usually regarded as a sample from some
population.
Factor analysis, on the other hand, seeks to account for, or "explain", the matrix of
correlations by a minimum, or at least a small number, of hypothetical variables or
factors. Factor analysis asks the question, "Does a random variable F 1 exist such that
the partial correlations between pairs of variables are zero after the effect of F 1 has
been removed?" If the correlation matrix is still unexplained, the question is asked
whether two random variables, F 1 and F exist so that the partial correlations between
2'
pairs of variables are zero after the effects of both of these variables have been removed,
etc. Thus, it may be said that principal components analysis is variance-oriented,
whereas factor analysis is covariance-oriented.
As has been noted above, the number of factors needed to explain the correlations is
fixed by the data itself, in the sense that when the factor extraction process leaves a
31
residual correlation matrix of approximately zero, all of the covariance present has
been accounted for.
However, this brings up the biggest problem in factor analysis. When a set of variables
is intercorrelated, we have a set of n(n-1)/2 correlations. This leaves unanswered
the question of what to put in the diagonal of the matrix, since we need a complete
matrix for the process of factor analysis. Two solutions to this problem exist: (1) put
ones in the diagonal — the method of principal components — on the grounds that,
except for errors in measurement, a variable should correlate perfectly with itself;
(2) insert values into the diagonal known as communalities. (the term communality
means the amount of variance of the variable accounted for by all the common factors
together). This will obviously be less than the total variance, since some of the
variance in any correlation matrix will be error variance, and some variance specific
to that variable. This second solution is called factor analysis.
Using the method of principal components, it is possible to account for many variables
by a few factors, since the first few principal components usually account for most of
the variance. However, unlike the factor analysis model, the variance accounted for
will include both specific factor and error variance. Factor analysis, in putting
communality estimates, instead of ones, in the diagonal, attempts to partition the
common factor variance from specific factor and error variance.
Unfortunately, use of the factor analysis model leaves the problem of deciding what
values to use for communalities. The communality of a variable is the most that it
has in common with other variables; thus, the squared multiple correlation of a variable
with all of the other variables constitutes a lower bound on the communality. The true
communality lies somewhere between the squared multiple correlation and one. To
date, no method has been found of arriving at the "true" communality.
The problem is further complicated by the fact that the communality estimates and the
number of factors extracted are mutually interdependent; the communality estimates
chosen and the number of factors chosen determine when the residual correlation matrix
drops to zero. Since, both of these values must be solved for simultaneously — and
this is impossible — one has to start with one fixed and allow the other to be decided
on by the program. Probably the best method to use is to fix the number of factors,
and by iteration find communalities that exactly fit the off-diagonals to give that number
of factors. In other words, decide on a number of factors, insert the squared multiple
correlations as initial estimates of the communalities, and factor the correlation matrix,
from which a new set of communalities is obtained; using these new values, refactor
the matrix again. This process is continued until the change in communalities between
successive factorings becomes trivial.
A final problem is deciding the number of factors to extract. In the absence of prior
knowledge of the number of common factors in the correlation matrix, the safest
course to adopt is Guttmann's lower bound theorem (1954), which demonstrates that
eigenvalues with roots less than 1.0 are statistically insignificant, so that one could
use this to set an upper bound on the number of factors to extract. Another possibility,
which can be used in conjunction with the method of principal components, is to extract
components until a prespecified amount of the total variance in the correlation matrix
has been extracted. This can be done most successfully when there is some idea as to
32
the amount of error variance in the correlation matrix (that is, the reliability of the
measurements is known). Thus, a slightly smaller proportion of the variance can be
extracted than is known to be common variance, since to take out more factors is to
include error variance.
In addition to finding out the number of factors (or components) required to account
adequately for an observed set of variables, we may also be interested in finding out,
or defining, what these variables are. When a predetermined number of factors are
extracted from a correlation matrix, we have a matrix of factor loadings, with k
columns (for the k factors), and n rows for the n original variables. These factor
loadings for variables v 1 , v 2 , ...vn on factors 1, 2, ...k are the correlations of the
newly discovered factors with the original variables.
The concept of simple structure applies to the factor loading matrix. As has been
pointed out above, the factor loading matrix consists of numbers such that each
number corresponds to a given factor and a given variable. A particular element in the
matrix indicates the extent to which that factor is represented in a given variable.
However, the particular configuration of numbers obtained in an unrotated factor
loading matrix is largely a function of the particular method used in extracting the
latent roots and vectors of the correlation matrix, and may have no empirical meaning.
The concept of simple structure was developed as a number of criteria for the
orthogonal rotation of the original factor loading matrix into such a position that the
factors extracted are readily identifiable in terms of the original variables. Simple
structure is a nonmathematical concept that sets up several criteria for rotation. The
first of these is the existence of a positive manifold. This means that — all other
things being equal — the factor loading matrix should have a minimum number of
negative values. But for many factor loading matrices corresponding to particular
correlation matrices, we may still have a very large number of factor loading matrices,
which equally well account for the intercorrelations, and which all have mostly
positive values. To further restrict the selection of the particular matrix, by which we
shall define the primary variables, we require that each column of the matrix shall have
a small number of high factor loadings, and a large number of near-zero loadings. We
also require that each row of the factor loading matrix shall have at least one near-zero
factor loading, and that at least one of the others shall be large-positive. In general,
the fewer the number of large loadings and the larger the number of near-zero loadings
in a row, the more simple the structure of that particular variable. The more variables
that can be expressed in the simplest form in terms of a relatively large number of
near-zero loadings, the more simple is the structure of the factor loading matrix.
However, the knowledge that there are a relatively large number of zeros in each row
and column of the matrix of factor loadings is insufficient; the relative positions of the
zeros and high loadings in the matrix must also be taken into account. For every pair
of factors, the factor loadings should be arranged so that not many variables have high
loadings on both. If variables have high loadings on the same factor, this implies that
they tend to measure the same factor. Also, for any pair of factors, a number of
variables should have high loadings on one factor and near-zero loadings on the other.
According to these criteria, if the loadings of one factor were to be plotted against
another, we should find that some variables would cluster about zero, some would be
33
high on one axis and low on the other, and vice versa for other variables. Thus, the
procedure for finding the best simple structure matrix for any given set of variables
would give an indication of which variables were the best measures for which factors.
The factors would then be defined in terms of the variables that have relatively high
loadings on them. In other words, simple structure is the application of Occam's
Razor to the factor loading matrix, since it aims to explain the configuration of
variables in such a way that each factor is represented by only a few variables (that is,
it loads or correlates with the smallest possible number of variables).
Practically, simple structure is realized by several computational techniques. The
best known is the Varimax rotation (Kaiser, 1958), which aims to maximize the fourth
power of the factor loadings; this amounts essentially to maximizing the scatter among
the loadings. Since a few highs means several lows, this leads to finding a position in
which there are many low loadings. Thus, Varimax aims to prevent a variable being
simultaneously highly loaded on two factors.
The Varimax rotation retains the property of orthogonality among the factors (that is,
the factors remain uncorrelated). In other words, the clusters of variables that load
highly on each of the factors are essentially uncorrelated with one another. In actual
practice, this seems rather unlikely. It would seem more plausible that the clusters of
variables (or factors), which are obtained from a factor analysis, are probably
correlated (positively or negatively) with one another. It is for this reason that oblique
simple structure rotations of the factor loading matrix are performed. The criteria
for oblique simple structure are identical to those for orthogonal simple structure, with
the exception that the orthogonality restriction is relaxed (that is, the factors need not
be uncorrelated with one another). Basically, rotation to oblique simple structure is
achieved by first performing a Varimax rotation, and then using the Varimax factor
loading matrix to rotate obliquely. The effects of oblique rotation are usually to
maximize the high loadings on each factor and minimize the near-zero loadings. The
clusters of variables on each factor, as derived from the Varimax rotation, will not be
altered; the effect simply is to give a "cleaner" solution.
When an oblique rotation is used, the situation is rather more complex than in the
orthogonal case. There are, in fact, three matrices that must be understood in relating
factors and variables:
1. The factor structure matrix, which gives the correlations between factors and
variables
2. The factor pattern matrix, which gives the loadings of factors on variables
3. The factor estimate matrix, which gives the beta weights for estimation of factors
from variables
The final feature is the estimation of factor scores, for both the orthogonal and the
oblique case. When a given number of factors have been extracted from a large number
of variables, and they have been identified by rotation of the factor loading matrix, it
may be desirable to calculate, for each observation, a set of factor scores, in place of
the scores on the original variables. These scores are a weighted summation of the
34
original scores on each of the variables for a given observation. These scores can
then give a profile score for each observation over the factors extracted.
2.2.1 Mathematics of Principal Components and Factor Analysis
Principal Components Analysis
Let us consider n points in a space of p dimensions, when the x's are expressed in
standard score form (that is, mean = 0; variance = 1). The line with current
coordinates X is
X 1 -m l
m
XP
P
X2-m2
-
g lg2
(1)
gP
where g's are direction cosines, and are subject to the condition
E g2
=1
(2)
i=1 .
The sum of squares of the distances from the n points on to this line is nS, where
n p
nS =
> (E
j=1 i=1
(x i• -mi )
2
- (
E gi (xi•-mi)) 2
i=1
(3)
If this is a stationary value, the partial derivatives with respect to the m's vanish,
and
-E (x..-m.) +Z gZg.(x..-m.) = 0, i = 1, 2, ..p
J
and since Z x. = 0,
j
.
(4)
i
= k (a constant).
Thus, the origin lies on the line (1), and we may take all the m's to be zero, thus
n 2
Lx.. = n,
j=1 "
and we have
n p 2
nS =
p
E ( E x.• - ( Egixii ) 2 )
j=1 i=1
i=1
n p
= np -
E
(E g.x..)
j=1 i=1 1
2
(5)
35
Then we can find stationary values of S for variations in g, subject to equation (2) above.
If, we consider X , an undetermined multiplier, this gives us
- l/n j=1
z
+ X gk = 0, k = 1, 2, .. . p
(6)
which gives us the set of p equations
g (1-X) + g r + ...gr = 0
1
2 12
pl
p
g
. r +g r +
2 p2
pl
+ g (1-X) = 0
(7)
Eliminating the g's, we get the characteristic equation of the correlation matrix
r-XII= 0
(8)
I
For a known r, this gives p roots in X. To each root corresponds a set of g's for
which S has a stationary value. Also, using equation (6), we find from equation (5)
S = p-
(9)
It follows from equation (9) that the root which gives the minimum S is the one with the
largest X . If we choose the largest root of equation (8), we have the line required. The
sum of squares of distances of the points from it is a minimum, and the variate
measured along it has the maximum variance. The variate is given by
V 1 = f.-/ g13.x.3
3=1
(10)
This indicates that the set of g's relates to X 1 . If we multiply, equation (6) by xk, and
sum over k, we can see that the variance of V 1 is Xi.
If we now look for the direction (perpendicular to the first line), for which the sum of
squares of perpendiculars is at a minimum, we find the line corresponding to X 2,
the second largest root, etc.
Thus, we have transformed to a new set of variates, V, that are uncorrelated, and
have variances X 1 , X 2 , ... X p in decreasing order.
Note that p = Z X, since the sum of the roots is the sum of p units in the main diagonal.
Also, if the original variables are normally distributed, we can regard the V's as
splitting off independent components of variance 1'
XX 2' ... X from the total
13'
variable, p.
36
Thus, we can select from this set of p variates, V, the first n components and consider
these as our factors. In general, the two main criteria for deciding on the number of
components to retain are (a) when the value of the X's falls below 1. 0; or (b) when the
first n components account for a fixed percentage of the total variance.
Factor Analysis Model
Basically, the mathematics and computational steps in factor analysis are similar to
those required by the method of principal components. The difference lies in the fact
that in component analysis we begin with a set of observations and look for components,
in the hope that we shall be able to effect a reduction in the dimensions of variation,
and that we can give some physical meaning to the components thus extracted. Using
the factor analytic model, we begin with a theoretical model, and try to find out whether
it agrees with the data, and if it does, to estimate its parameters.
Let us begin, as in the method of principal components, .with a matrix of observations
and consider whether they can arise from a situation with the following structure:
p
1
7x=
a f +bs +ce.,
1.-4 ik k
i i
i
k=1
where i = 1...p
(1)
In this equation (1), the f k are factors that can appear in more than one x, s i is a
factor specific to the variable x i , and ei is an error term.
At this level, the model is undetermined, and, in fact, by using the method of principal
components, we can always express the x's in terms of f's without invoking specific or
error terms at all. This is the basic difference between principal components and
factor analysis. In the former, we consider all the variance, common variance,
and extract its orthogonal components. In factor analysis, on the other hand, we take
into account that some of the variance is going to be due to error, and some to variance
that is quite specific to a certain variable. In this sense, the factor analytic model is
more realistic, and the principal components method, in spite of its mathematical
simplicity is misleading.
Therefore, let us assume that we have n observations on p variables xi . Using the
method of principal components, we have
S = E var xi - g ii gikcov (xj , xk)
(2)
Thus, the principal components equations are
z.cov(x x) =X . .
.i j
j
kk
(3)
37
However, in the factor analysis model, if
x.. =Za. f
kj
k
aikaml/nZfkjfmj
cov(x.,x.) =
13
k, m
(4)
and if we substitute its expected value for E f .1 .
ki mi
if k = m
cikm =
=0,
if k m
we have
cov(x.,x ) = a. a.
k
(5)
x. =Z a. f + e.
kj
(6)
cov(x., x.) =Z a. a. + var e.d.
j
jk
(7)
If the model is
we have
Now we are required to estimate the coefficient
matrix
a
I
We can operate on the estimated
a + d var e.
ik k
i
This is the same as the former matrix, except for the principal diagonals, where each
term is increased by var e i . Thus we would like to have in the main diagonal not
Eat + var e., but only Z a2ik. In other words, if we are not to bias the estimates of the
ik
a's, we must remove var e from the diagonal terms. This is equivalent to substituting
communalities for unity in the diagonals of the standardized matrix.
Possibly the best way of estimating the communalities is to start with the squared
multiple correlations in the diagonal (since this is a lower bound on the communalities).
We then perform an analysis of the data and arrive at certain factors (deciding on the
number of factors, by taking those with latent roots >1. 0, those factors accounting for
a certain percentage of the variance, or some other method). We can then use the
coefficients occurring in those factors to estimate new communalities, iterate, and
proceed until the communalities converge. Basically, what this process amounts to is
that we assume m factors and assume that they account for as much as possible of the
38
variance; this determines the communalities and, consequently, the "error" variances.
But this does not mean that we have estimated the actual error variances that occur in
practice. We have estimated only what they would be if the number of factors is what
we think it is, and the error variances are minimal.
In other words, this would suggest that in practice care should be exercised in
computing a factor analysis. If one has no idea of the number of factors to extract,
possibly the best solution is to compute a principal components solution of various
numbers of factors, rotate, and decide which set of factors gives the best empirical
meaning. Then, using this number of factors, estimate the communalities, and run
a proper factor analysis.
Rotation of the Factor Loadings
In using the Varimax criterion for rotation to orthogonal simple structure, we have
already defined verbally what we mean by simple structure. Mathematically, the
simplicity of the factorial composition of the j th variable can be defined as the variance
of the squared loadings for the test
2 2 2
22
cr.i` = (rZ (a. ) - (Z a.
)/r
is
is)
3
s
s
(1)
where j = 1, 2, ... n variables, and s = 1, 2, ... r factors, and a is is the factor
loading of the j th variable on the s th factor.
To obtain the total criterion for the entire factor matrix, we can sum over the variables
thus:
2 2
2
q* = Z (rZ (a. ) 2 - a a.
)/r2
is
is )
s
j
s
(2)
This criterion can be modified if we define the simplicity of a factor as the variance
of its squared loadings
v* = (n Z (a. )
s
.
i2s
3
2 - ( Z a.2 )2)/n2
. is
3
(3)
and for the criterion for all the factors, define the maximum simplicity of a factor
matrix as the maximization of
2
v* = Z v* = Z((n Z (a.is )2 )/n2
s
s
j
which is the variance of squared loadings by columns rather than by rows.
This is the row Varimax rotation. However, this will exhibit a systematic bias
because of the divergent weights that implicitly are attached to the variables by their
39
communalities. Therefore, the normalized Varimax rotation weights each variable by
its communality, thus:
2
2 2 2
2 2
v = I ((nZ (a. /h. ) 2 - (z (a. /h. )) )/n )
js j
js j
s
(4)
2
jth
where is the communality of the — variable. In this case, the variance of the
11J
squared correlations of the common parts of the variables with a factor are now being
maximized.
The oblique rotational scheme is developed from the above normalized Varimax
solution. The Promax rotation simply takes the Varimax-rotated factor loading matrix
and generates a pattern matrix from it by powering all the elements in the original
matrix.
(p..)
We can define a matrix P..
such that
k+1
..= a..
pij
/ a..
(5)
with k > 1. Each element of this matrix is, except for the sign that remains unchanged,
the ktil power of the corresponding element in the row-column normalized orthogonal
matrix. We then find the least-squares fit of the orthogonal matrix of factor loadings
to the pattern matrix generated by equation (5).
L = (G'G) 1G'P
(6)
where L = the unnormalized transformation matrix of the reference vector structure,
G = the orthogonal-rotated matrix, and P = the matrix derived by equation (5) above.
Finally, the columns of L are normalized so that their sums of squares are equal to
unity.
Computation of Factor Scores (Direct Estimation)
Factor scores can be computed for each observation on the factors extracted from the
original correlation matrix. (See Harman, 1960, Chapter 16, in list of references at
end of this chapter.) These are computed using the following equations:
Let F = matrix of factor loadings
V = the orthonormal matrix of eigenvectors
A = diagonal matrix of latent roots (eigenvalues)
Then
F = VA
40
1/2
(1)
F
is an orthogonal matrix (but not orthonormal). Since V is orthonormal
V' = V
-1
(2)
From the principal components model we know
Z = BS + US + eS
(3)
where Z = vector of scores for each observation
B = matrix of common factor loadings
U = diagonal matrix of uniqueness (1-h2)
e = diagonal matrix of errors
Therefore
S = F -1 Z
(4)
since
F
-1
-1 -1/2
=V A
= VA
-1/2
(5)
SO
-1
-1
= A F'
S = A 1F'Z
F
(6)
(7)
Therefore, the factor score for an observation on factor 1 would be
s1=f
Z1
/X
Z
11
1 +f
12Z2/X
1 +...+f1p
pA 1
(8)
TERMS USED IN FACTOR ANALYSIS
Communality. Sum of squares of factor loadings for any given variable (that is, the
total variance due to factors which this variable shares with other variables in the
matrix).
Covariance. Mean
(1/n) Z (x-R) (y-57).
product of deviations of variable x and variable y from their means
41
Diagonal matrix. A square matrix having zeros in all positions except those on the
diagonal from upper left to lower right.
Direction cosine. One of a set of cosines of angles, defined for a point, each angle
being measured between one of the reference axes and the vector connecting the point
with the origin.
Factor loading. Correlation of any particular variable with the factor being extracted.
Factor matrix. Matrix whose entries are the factor loadings obtained from a factor
analysis; it generally is arranged so that it has as many columns as factors extracted,
and as many rows as variables.
Hyperplane. Space of (N-1) dimensions, defined by a reference vector perpendicular
to it. (For example, in two dimensions, either coordinate axis is the hyperplane of
the other; in three dimensions, the plane defined by any two coordinate axes is the
hyperplane of the third.)
Normalize. To divide each of a set of numbers by the square root of the sum of squares
of all numbers in the set, so that the sum of squares of the new set is 1.00.
2.2.2 The Program
Given a set of observations numbering 499 or fewer containing measures on n 530
variables, x 1 , x2 , ...xn , a square symmetric n x n correlation matrix, R, and vectors
of means and standard deviations are computed. From this or a given correlation
matrix, a factor matrix is extracted containg n or fewer vectors of factor coefficients.
This factor matrix may be rotated to approximate simple structure by an analytical
criterion in either an orthogonal or an oblique reference frame. Given the original
data, X, together with its means and standard deviations, this rotated factor matrix, G,
may be used to compute factor measurements (factor scores) by a regression model.
Communalities, h?, are treated as the diagonal elements of the correlation matrix.
These elements are computed equal to 1.0, and should be retained as such for the
computation of the principal components factor matrix. They may be specified by some
predetermined values, however, where a specially constructed correlation matrix
is given rather than carried over from previous computation. For example, the user
may desire to place specific communality values on the diagonal of this matrix (see
section 2.2.6). Options are also available to consider each 1--.thcommunality as the
maximum absolute off-diagonal element in the .th vector of the correlation matrix or
.th variable.
as the squared multiple correlation for the 1-All n of the latent roots of the correlation matrix with diagonal elements chosen as
indicated above are computed by a Householder (HOW, 1962) tridiagonalization, followed
by the use of the QR algorithm. The k latent vectors are computed by Wilkinson's
(HOW, 1962) method. These latent roots (eigenvalues) are solutions to the matrix
equation
(R-X) A* = 0
where 11 is the correlation matrix, X is a diagonal matrix of the latent roots, and A*
is a matrix of the latent vectors (eigenvectors).
42
a
The total variance accounted for by the principal components of the correlation matrix
is evaluated by its trace (the sum of its diagonal elements). The percentage of this
total variance accounted for successively by each latent root is computed and presented
cumulatively for the first through the nth root. This percentage is presented whether
1, the
or not a principal components analysis (rii = 1) is being performed. If r ii
output should be ignored. The latent vectors, a'!`, are normalized, and a matrix, F, of
factor loadings is computed by scaling each latent vector by the square root of its
associated latent root, that is,
where the 1.. are factor loadings, the a..
are elements of the latent vectors, and the
1.)
X. are latent roots.
Several methods are available to the user for estimating the rank, m, of the factor
space for the purpose of retaining only m factors for output or subsequent rotation. The
value of m may be specified on a control card and arbitrarily accepted as the maximum
number of factors. The user may also request that only those factors whose latent
roots are equal to or greater than 1.0 be retained. An option is also available to retain
only those factors which cumulatively account for an amount of variance equal to or less
than a given percentage of the total variance. This option could be used in a principal
components analysis.
The matrix of factor loadings may be rotated to approximate simple structure in an
orthogonal reference frame by the Normal Varimax method (Kaiser, 1958) for the case
of uncorrelated factors; k factors are rotated where k is equal to or less than m, the
rank of the factor space as determined above. The Normal Varimax method develops
a transformation matrix, T, over a cycle of rotations of each of the 2(k-1) pairs of
orthogonal axes of the factor space taken in turn. The angle of each rotation is
chosen such that a function, U, of the factor matrix is maximized. Complete cycles
of rotations are performed until U is not significantly increased by an additional cycle.
U is computed by an evaluation of the following expression:
n k [g]
u = kE E
i=1 j=1
4
k[n
- E E
j=1 i=1 h
22
2
where k is the number of factors, n the number of variables, and glj
element of the
factor matrix under rotation for the ith variable on the j th factor; h2represents the
communality of the .th
1-- variable computed using only the k factors under rotation. The
final rotated factor matrix, G, is derived by the matrix multiplication, G = FT, where
T is the complete Varimax transformation matrix, and F is the unrotated factor matrix.
This matrix of k factors may also be rotated to approximate simple structure in an
oblique reference frame by the Promax method (Hendrickson and White). A pattern
43
matrix, P'; describing a factor matrix rotated to approximate simple structure in
oblique axes, may be accurately estimated by a matrix whose elements are functions
of the elements of the orthogonal matrix rotated by the Varimax method. This matrix
can be derived by the following operation:
a+1
g• .
P*
ij
gi]
where P fl`• is an element of the pattern matrix, P:I` and g. • is an element of the
orthogonally rotated factor matrix, G. The value of a is four (4). According to
Hendrickson and White (1964), four is the optimal value in the majority of cases.
The user can, however, easily change this number in the program PROMX. In general,
as a increases, the dependence of the rotated factors or their obliquity increases. A
transformation matrix is computed that rotates the orthogonal factor matrix into an
oblique reference vector structure matrix, V, which is a least-squares fit to the
pattern matrix, P*, described above. This transformation matrix, L, is derived in
unnormalized form from the following matrix equation:
*
L = (G'G) 1G'P
After the transformation matrix, L, has been column-normalized, the reference
vector structure matrix, V, is obtained from V = GL. The correlation matrix of the
reference vectors, lk , is computed by 4, = L' L, and the reference vector pattern matrix,
W, is then developed by W = VI -1 . To derive the primary factor structure from the
reference vector solution described above, the diagonal matrix, D, of the correlations
among the reference vectors and the primary factors is computed by taking
where the d.. are diagonal elements of D, and c., are diagonal elements of the matrix,
1tt
vi . The primary factor structure matrix, S, is then determined by S = WD, and the
primary factor pattern matrix is derived by P = VD- 1 . The matrix of correlations
between the primary factors, 43 , is computed by ct. = DI -1D.
Factor measurements (factor scores) are computed by the "short" regression method
(Harman, 1960). The diagonal matrix of uniquenesses, U, is obtained by taking
2
2
1-h . where the h. are the communalities computed from G, the orthogonal factor
u
matrix, or P, the oblique factor matrix, depending on which scores are requested. For
the case of uncorrelated or orthogonal factors, the matrix, Q, is developed by
-1
Q = I + G'U G, where G is the orthogonal factor matrix (Varimax solution). Factor
scores are formed by the operation f
= 13' Z,
where
7
is a factor score matrix,
(3
is a
matrix of factor score regression coefficients, and Z the vector of standardized data.
ft is obtained by the operation
44
=
Q-1G 1 U-1 .
The elements of Z are computed by
x .-x.
th
zki = IV, where x,,c is an observation in a data matrix for the k — sample case on
th
the j11-.1 variable, and xj is the mean and s j the standard deviation of the j — variable.
For the case of oblique or correlated factors, the procedure is much the same, except
-1
1
for the definition Q = 4)
+ P'U P, where P is the primary factor pattern matrix and
1
-1
4) the matrix of intercorrelations of the primary factors. Also, here, j3' = Q P'U .
2.2.3 Summary of Output Statistics
1. High and low value of each variable
2. Means of each variable
3. Standard deviation of each variable
4. Sample variance for each variable
5. Matrix of raw cross products
6. Matrix of residual cross products
7. Variance-covariance matrix
8. Matrix of correlation coefficients
9. Matrix of characteristic vectors
10. Characteristic values
11. Trace
12. Cumulative percentage of trace of each characteristic value
13. Unrotated factor matrix
14. Orthogonal transformation matrix
15. Orthogonal factor matrix
16. Transformation matrix to oblique reference vector structure
17. Oblique reference vector structure matrix
18. Correlations among oblique reference vectors
19. Oblique reference vector pattern matrix
20. Oblique primary factor structure matrix
21. Correlations among oblique primary factors
22. Oblique primary factor pattern matrix
23. Factor score regression coefficients
24. Factor scores
25. Communalities
45
2.2.4 Job Execution
To perform a factor or principal components analysis, the user must supply three sets
of cards to the program:
1. Monitor control cards
2. Program control cards
3. Data cards
Descriptions of the form and content of each card set follow.
MONITOR CONTROL CARDS
The monitor control cards are necessary to initiate program loading from the disk and
to establish the necessary communication with the monitor. A general description of
the cards may be found in IBM1130 Disk Monitor Reference Manual (C26-3750).
A factor analysis requires the following monitor cards:
CC: 1 4
8
1/ XEQ FCTR
16-17
04
*LOCALFCTR, FMTRD, DATRD, PRNTB, MXRAD, TRAN
*LOCALFCTR1, TRIDI, QR, INVRS
*LOCALFCTR2, VECTR, PRNT
*LOCALFCTR3, VARMX, PROMX, SCORE, RFOUT
*LOCALCOREL, PRNT
The monitor control cards do not change from job to job within one analysis, but must
be included with every job processed. The first program operated on by this system
should be preceded by a cold start card.
PROGRAM CONTROL CARDS
The program control cards communicate the data-specific parameters and output
options to the program. There are five possible card types necessary for execution:
1. Input/output units card*
2. Job-title card*
3. Option card (described below)
4. Variable name card (described below)
5. Variable format card*
Four of the control cards are required in every job. The variable format card is
necessary only if source data is to be processed.
*See "General Operating Instructions", section 1.2.
46
OPTION CARD
Number of Variables (cc 1-2)
This field must be punched with a nonzero integer, n, which is less than or equal to 30.
The value contained in n is the total number of variables to be processed.
Input Type and Source (cc 3-4)
This field allows the user to specify the input device (1442 card reader or disk) and,
indirectly, the type of input analysis to be undertaken in the input program. The three
possible values that may be punched in this field are described below:
Meaning
Value
1
Raw data will be read from the 1442 card reader and transferred
to the disk. It will be retained there for use by this or other
programs until destroyed by input from one of the four programs
in this system. Raw sums and raw sums of cross products will be
accumulated. Data will be read until a card with a negative number
in the identification field is encountered (section 2.2.5).
2
Raw data will be read from the disk. Raw sums and raw sums of
cross products will be accumulated. Data will be read until a
negative integer in the identification field is encountered.
3
A previously computed matrix, or matrices, will be read from
the 1442 card reader. Matrix cards will be read until a negative
job number field is encountered (see "Pooling", section 2.2.6).
Sequence Checking (cc 5-6)
This field is used to indicate that raw data input from the card reader (cc 3-4 contains a
1) is to be sequence-checked. A value of zero or a blank field implies that no sequence
check will be made. A value of one (1) implies that the cards will be sequence-checked.
The sequence-checking process consists of an equal comparison check of the case
identification field for all cards in a case and an ascending sequence check of the card
number field. If an error in either of these conditions is encountered, the program
prints a message, and the job is terminated.
Number of Variables on Card 1 (cc 7-8)
When a data vector contains more variables than will fit on one card, the user must
indicate to the program the number of variables punched on each card. This field must
be punched with the number of variables on the first card. If there is only one card
per case, this field must be blank or zero.
Number of Variables on Card 2 (cc 9-10)
Same as cc 7-8, except that this field indicates the number of variables on the second
card of the data.
47
Number of Variables on Card 3 (cc 11-12)
Same as cc 7-8, except that this field indicates the number of variables on the third
card of the data.
Transformation Switch (cc 13-14)
If the value in this field is nonzero, a user-written transformation subroutine is called
after each data record is read and before any computation takes place.
If the value in this field is zero or blank, the transformation subroutine is not called.
The use of transformations is discussed in section 2.5.1.
Output Raw Sums of Cross Products (cc 15-16)
This field is used to indicate whether the raw sums and sums of raw cross products
matrix are to be printed, punched, printed and punched, or not presented.
The four (4) possible values of this field are described below. The computation to
generate the matrix is performed even if the "no output" option is chosen.
Value
Meaning
0 or blank
1
2
3
No output.
Matrix will be printed.
Matrix will be printed and punched.
Matrix will be punched.
Punched , output of the raw sums of cross products matrix includes the number of
observations and the vector of raw sums and sums of squares. This entire output
must be entered on the pooling option (section 2.2.6).
Output Residual Cross Products (cc 17-18)
This field is used to indicate whether the residual cross products matrix — defined
as:
u..
1)
= c..
13
S .S.
I
n
i,j = 1,2...n
where c.. are the elements of sums of raw cross products matrix, and s., s. are the
I 3
13
th
th
raw sums of the i — and j — variables, respectively, and n is the number of cases — is
to be printed, punched, printed and punched, or not presented.
The four (4) possible values are described above under "Output Raw Sums of Cross
Products".
The matrix is computed even if the "no output" option is chosen.
48
Output Variance-Covariance Matrix (cc 19-20)
This field is used to indicate whether the variance-covariance matrix — defined as:
u..
c.. =
1.)
i,j = 1,2...n
n-1
where u ij. • is an element of the residual cross products matrix, and n is the number of
cases — is to be printed, printed and punched, punched, or not presented. The four
(4) possible values that may occur in this field are as is given for the above matrices.
There are no additional vectors or matrices punched with the punched output. The
matrix is computed even if the "no output" option is chosen.
Output Correlation Matrix (cc 21-22)
This field is used to indicate whether the correlation matrix — defined by:
i,j = 1,2...n
where c.. is an element of the variance-covariance matrix, and s., s. are the standard
j
th
th
deviations of the i — and j— variables, respectively — is to be printed, punched,
printed and punched, or not presented. The four (4) possible values contained in this
field are as is given for other matrices, above.
The punched output of the correlation matrix includes the number of cases and cards
containing the vectors of means and standard deviations.
The matrix is generated even if the "no output" option is chosen.
Factor Scores (cc 23-24)
This field is used to indicate whether factor scores are to be computed. If a value of
zero (0) is punched or the field is left blank, the factor score computation is suppressed.
When this field contains a one (1), factor scores and factor score regression coefficients
are computed. A two (2) in this field causes scores to be punched. Upon entry to the
program SCORE, which computes the scores, the program assumes that the necessary
matrices and data have been set up for the analysis. Hence, either an orthogonal
or an oblique rotation must be performed before entry to the SCORE routine. The
number of scores to compute is equivalent to the number of rotated factors. Hence, it
is not possible to compute factor scores from the unrotated principal axis factor matrix.
In addition to the factor matrix and an auxiliary matrix computed by the rotation output
program, the scores program requires the data file to be on the disk and the means and
standard deviations vectors to be located in common storage. If the entire factor
analysis is being done from the raw data, these operations are performed automatically
by the program.
49
Factor Score Punched Output Format:
Card I
Columns
1-4
5-6
7-20
Elements
21-34
Factor score number (I)
25 (Identifier)
Score for variable 1; ± 0 .XXVOCXIXE±XX
Score for variable 2;
•
63-76
Score for variable 5;
II
If more than five factors have been rotated, card I + 1 has a format as follows:
Columns
1-6
7-20
Elements
Blank
Score for variable 6
The factor scores are computed from the Varimax solution if an oblique rotation has
not been requested; otherwise, they are computed from the oblique solution.
Number of Factors to Compute (cc 25-26)
The number punched in this field is used as a switch setting in the program to determine
the number of factors to compute from the characteristic roots and vectors. There are
four (4) possible values that may be punched, and they are described below:
Value
50
Meaning
0 or blank
No factors will be computed.
1
Only those factors will be computed whose characteristic
vectors have associated characteristic roots greater than or
equal to one (1).
2
Compute a fixed number of factors (m). The value of m will
appear in cc 27-28; m must not be greater than ten.
3
Compute factors whose variance accounts jointly for no more
than P percent of the total variance. The value of P appears
in cc 27-28. The variance of a factor is the characteristic
root associated with a particular characteristic vector. The
Meaning
Value
3 (cont)
total variance is defined as the trace of the matrix. The
percentage is computed by adding characteristic roots and
forming the ratio of this sum to the trace of the communalityadjusted correlation matrix.
Constant for Number of Factors (cc 27-28)
This field is used in conjunction with cc 25-26. A two (2) in the previous field implies
that this field will contain an integer, m, which is equal to the number of factors to
compute. If cc 25-26 contains a three (3), this field should contain an integer, P,
which is the percentage of factor variance.
Communality Estimation Options (cc 29-30)
This field is used to offer the user a choice of three methods of estimating the
communality or common variance of the reduced factor space. Before the characteristic
roots and vectors are computed, the program places the communality estimate on the
principal diagonal of the matrix to be factored. Three possible values may be punched
in this field, and they are described below. Iteration on communalities is not performed
automatically, and is discussed in section 2.2.6.
Value
Meaning
0
No change to the matrix.
1
The absolute value of the largest off-diagonal element in a row
will be used as the communality estimate for that variable.
2
The square of the multiple correlation coefficient between variable
i and all other variables in the matrix will be used as the
communality estimate for the
variable. This is done for all i.
Rotation Switch (cc 31-32)
This field is used to indicate the type of rotation to simple structure that will be used
on the principal axis factor matrix. The user has the choice of choosing an orthogonal
rotation (normal Varimax) and/or an oblique rotation (Promax). The process by which
an oblique rotation is computed requires that an orthogonal rotation be performed first.
Hence, in addition to the oblique rotation matrices, the user has the option of obtaining
the output of the orthogonal rotation. Three possible values may be punched in this field,
and they are described below:
Value
0
1
2
or blank
Meaning
No rotation will be performed.
Orthogonal rotation.
Oblique rotation (includes an orthogonal rotation).
51
Number of Factors to Rotate (cc 33-34)
This field is used in conjunction with the rotation switch described in the previous field.
The value punched in this field determines the number of factors to rotate. Two
possible conditions can arise. If the user does not know the number of factors to
rotate, it is suggested that the field be left blank (or zero). The number of factors to
rotate is then chosen on the basis of one of the options in cc 25-26. However, if a
value of k appears, k factors are rotated if this number is less than or equal to the
number of factors computed. In any case, k must be less than or equal to ten.
Pooling Option (cc 35-36)
When using the matrix input/output option (03 in cc 3-4) and when pooling sums of squares
and cross products (section 2.1.4), if the user desires that matrices be subtracted
rather than added, this field should be nonzero.
Factor Matrix Output Option (cc 37-62)
In a complete factor analysis, there are 13 additional matrices (section 2.5.3) that the
user has the option to output, if desired. The 13 remaining fields on the card are for
this purpose.
Each of the two-column fields may take four possible values, described below:
Value
Meaning
0 or blank
1
2
3
No output.
Print matrix.
Print and punch matrix.
Punch matrix.
The following gives the name and field column numbers for each matrix:
Column
37-38
39-40
41-42
43-44
45-46
47-48
49-50
51-52
53-54
55-56
57-58
59-60
61-62
52
Matrix
Characteristic vectors, A*
Unrotated factors, F
Orthogonal transformations, T
Orthogonal factors, G
Transformation to oblique reference vector structure, L
Oblique reference vector structure, V
Correlations among oblique reference vectors, ‘G
Oblique reference vector pattern, W
Correlations between reference vectors and primary
factors, D
Oblique primary factor structure, S
Correlations among oblique primary factors, cl)
Oblique primary factor pattern, P
Factor score regression coefficients,
Factor Analysis Option Card Summary
Column
Meaning
1-2
Number of variables
3-4
Input type and source
1 - Raw data input from card reader
2 - Raw data input from disk
3 - Matrix input from card reader
5-6
Check sequence of raw data input
0 - No
1 - Yes
7-8
Number of variables on card 1
9-10
Number of variables on card 2
11-12
Number of variables on card 3
13-14
Transformation switch
0 - No transformation
1 - Transformation
15-16
*Output raw cross products matrix
0 - No
1 - Print
2 - Print and punch
3 - Punch
17-18
*Output adjusted cross products matrix
0 - No
1 - Print
2 - Print and punch
3 - Punch
19-20
*Output variance-covariance matrix
0 - No
1 - Print
2 - Print and punch
3 - Punch
21-22
*Output correlation matrix
0 - No
1 - Print
2 - Print and punch
3 - Punch
* Not available when correlation matrix is used as input.
53
Column
54
Meaning
23-24
Factor score options
0 - Do not compute factor scores.
1 - Compute and print factor scores.
2 - Compute, print, and punch factor scores.
25-26
Number of factors options
0 - Do not compute factors.
1 - Compute factors for latent roots 1.0 only.
2 - Compute m factors (where m is given in cc 27-28).
3 - Compute factors accounting jointly for no more
than p percent of the total variance (where p is
given in cc 27-28).
27-28
Constant for number of factors option, if appropriate
29-30
Communality options
0 - Use diagonal values of correlation matrix
(normally unity unless otherwise specified in
a given matrix).
1 - Use maximum absolute off-diagonal element
in each vector of the correlation matrix.
2 - Use the squared multiple correlation
coefficient for each variable.
31-32
Rotation options
0 - Do not perform any rotations.
1 - Perform an orthogonal rotation (Varimax) only.
2 - Perform an oblique rotation (Promax,
including Varimax).
33-34
Constant for number of factors to rotate
0 - Rotate the number of factors determined by
the option chosen in cc 25-26 above.
k - Rotate a number of factors equal to the minimum
of k, ten, and/or the number of factors
determined by the option above in cc 25-26.
35-36
Pooling option (see Sections 2.5.3 and 2.2.4)
00 - Add matrices with ID = 1
Nonzero - Subtract matrices with ID = 1
Column
Meaning
Note: In columns 37-60, the matrix output options are as follows:
0 - No output
1 - Print only
2 - Print and punch
3 - Punch
37-38
Output the latent vectors, A*
39-40
Output the unrotated factor matrix, F
41-42
Output the orthogonal transformation matrix, T
43-44
Output the orthogonal factor matrix, G
45-46
Output the transformation matrix to oblique
vector structure, L
47-48
Output the oblique reference vector structure
matrix, V
49-50
Output the correlations among oblique reference
vectors,
51-52
Output the oblique reference vector pattern
matrix, W
53-54
Output the correlations between reference vectors and
primary factors, D
55-56
Output the oblique primary factor structure
matrix, S
57-58
Output the correlations among oblique primary
factors,43
59-60
Output the oblique primary factor pattern matrix, P
61-62
Output the factor score regression coefficients, $
VARIABLE NAME CARD
In the factor analysis program there are a number of matrix printouts that the user may
request. The variables in the matrix may be assigned a four-character name to aid in
the identification of the output. The card is punched in four-column fields, and each
field corresponds to the variable to be identified (for example, field 3 (columns 9-12)
will be the name of row and column 3 on all matrix output). At most, 20 names can
appear on one card. If there are more than 20 variables in the analysis, a second card
having the same format as the first must be included in the control card deck.
55
Column
Meaning
1-4
Name of variable 1.
5-8
Name of variable 2.
(4N-3) - (4N)
Name of variable N.
2. 2. 5 Data Input
Raw data input to the program consists of a set of observations made on several
different variables. The variables for each observation are punched on one, two, or
three cards, according to the following general format:
Meaning
Field
Type
1
Integer (I) Identification field. Any numeric information that
serves to identify the particular observation is
punched in this field. It must be greater than zero,
and should be different for each observation.
2
Integer (I) Card number within observation. If it is not possible
for one card to contain all the variables, they may
be continued on a second and a third card, as
necessary. The user has the option of sequence
checking the cards to ensure that all cards within a
case are together, and that the order of cards is
consistent. If the option is chosen (cc 5-6 on option
card), this field must be punched with an integer
that is in ascending sequence for all cards in the
case. If sequence checking is not desired, the field
may be blank and may consist of one blank column.
3, 4. . . . , n
etc.
Floating
point (F)
Variable x l . Any number may be punched in this
field. Decimal points are not required.
The remaining fields on the card are reserved for
variables. If there are more variables than can
fit on the first card, a second and a third card may
be used.
Following the data deck, the user must include a card containing a negative integer
in the identification field. This card signals the end of data.
56
2.2.6 Matrix Input/Output
It is possible to obtain punched card output of a number of matrices (see section 2.5.3)
and vectors with this program. This program is designed to also input some of these
matrices, at a later time, for further analysis or processing. In addition, matrices
from another program or source, if punched in the program format, may also be used as
input.
This section is devoted to a description of various possible forms of analysis with the
output options available in each program.
Format Description
See the matrix format given under "Format Description" in section 2.1.4.
Factor Analysis with Correlation Matrix Input
The punched output option of the correlation matrix includes the punchout of the
number of cases (matrix 21), and means and standard deviation vectors (matrix 23).
This complete output can be used as input to initiate another analysis without the
necessity of reprocessing the source data that was used to generate the matrices.
To use the correlation matrix set as input, the user places the punched output behind
the variable names card, followed by a blank card, or one that contains a negative
number in the job number field. The program reads the number of cases, means, and
standard deviations and correlation matrix, stores each in its appropriate location, and
then initiates the analysis as specified on the option card.
Pooling Sums of Squares and Cross Products (cc 35-36)
The topic is discussed fully in section 2.1.4. When raw sums of squares matrices have
been previously punched by this program (or by hand), in accordance with the above
format description, they can be stacked and will be combined by using this option.
When the option is given the number 0, they will be added. If a nonzero field is used,
all matrices will be added until the first negative (left-justified) job number field
(cc 1-4) is encountered. Subsequent matrices will be subtracted until the second
negative problem number card is encountered.
Iterating on Communalities*
This program does not iterate on communalities, which is a desirable feature mentioned
in section 2.2. However, by electing to punch the correlation matrix, one can insert the
estimated communalities on the diagonal of this matrix and iterate using matrix input.
*The factor scores computation reads the original data matrix, X, from the disk. If
X is not on the disk, which is the case if any card reader input with other data has
been read subsequent to the reading of X, then X must be reread under input mode 1
before this matrix input option is used.
57
2. 2. 7 Operating Instructions
A.
Using the factor analysis program when the total 1130 Statistical System has not
been stored on the disk
If the user wishes to load only the set of programs that allow this type of analysis, the
following programs must be compiled or assembled and stored on the disk. Each deck
begins with a card punched as
// FOR
and ends with an
*STORE
card.
The user should use a disk containing the 1130 Disk Monitor System, as described in
section 1.1. The following decks should be preceded by a cold start card, placed in the
card reader hopper, and the buttons IMMEDIATE STOP (console), RESET (console),
START (card reader), and PROGRAM LOAD (console) should be pressed. A blank card
should be placed after the last deck in the card reader hopper.
DECKS-LABELS: FCTR-FCTR; FCTR1-FCT1; FCTR2-FCT2;
FCTR3-FCT3; *FMTRD-FMRD; *DATRD-DTRD; *GMPYX-GMPY;
*GDIVX-GDIV; *PRNTB-PRNB; **COREL-CORL; **PRNT-PRNT;
**1VIXRAD-MXRD; INVRS-INVS; XMAX-XMAX; TRIDI-TRID; QR-QR;
VECTR-VC TR ; COVEC -CVEC ; RFOUT- ROUT ; PROMX-PRMX;
VARMX-VR1VDC; RPRNT-RPNT; MATIN-MATN; SCORE-SCOR;
*FMAT-FMAT; TRAN-TRAN.
In addition, regression and factor analysis programs must reside on the disk together;
section 2.1.5 names additional routines to be placed on the disk.
B. Execution from disk
Once the component subroutines and main calling programs are on the disk, the
execution of a job requires the monitor control cards, program control cards, and data
cards to be placed in the card reader. The deck should be preceded by a cold start card.
To initiate processing, the buttons IMMEDIATE STOP and RESET (console), START
(card reader), and PROGRAM LOAD (console) should be pressed. The order in which
the cards are placed in the card reader for either matrix or raw data input is shown in
Figures 8, 9, and 10.
*Used in all four analysis types
**Used in regression analysis
58
(Optional Blank'I
Output
/ Negative
Identification
Data Deck
Variable
( Format
Variable
( Names
(Option
(Job Title
(Input/Output
Units
I
y
—
Monitor
Control
Figure 8. Factor analysis card order — card reader input
Optional
Blank Output
/Variable
Names
(Option
(Job-Title
(Input/Output
Units
Monitor
Control
Figure 9. Factor analysis card order — disk input
Optional Blank
( Output
/Negative
Identification
(Matrix
Matrix
/Variable
Names
/ <—(to be added to matrix 1)
/
(Option
(Job-Title
(I/0 Units
/Monitor
Control
Figure 10. Factor analysis card order — matrix input
59
2.2.8 Sample Problem
INPUT
05
// XEQ FCTR
*LOCALFCTR,FMTRO,DATRO,PRNTB,MXRAD,TRAN
*LOCALFCTR1,TRIDI,DR,INVRS
*LOCALCOREL,PRNT
*LOCALFCTR2,VECTR,PRNT
*LOCALFCTR3,VARMX,PROMX,SCORE,RFOUT
020200
FACTOR ANALYSIS SAMPLE PROBLEM
3333
040100000000000101010202020200020000 1010101010101010101010101
P1
P2
P3
P4
1212,1X,4F6.01
0101 000063000075000159000041
0201 000101000092000142000049
0301 000119000098000131000068
0401 000157000101000124000092
0501 000178000104000119000097
0601 000147000106000118000102
0701 000128000108000116000109
0801 000113000107000116000066
0901 000094000107000115000044
1001 000111000104000117000069
1101 000139000110000104000117
1201 000157000107000100000118
1301 000169000111000075000157
1401 000145000109000079000107
1501 000079000095000096000064
1601 000049000086000111000047
1701 000048000077000111000032
1801 000041000069000106000022
1901 000066000062000097000017
2001 000111000074000092000045
2101 000164000104000088000097
2201 000170000117000039000164
2301 00020800013 5 000053000246
2401 000237000148000058000366
2501 000169000152000061000230
2601 000114000137000073000175
2701 000106000130000077000178
2801 000097000123000086000156
2901 000099000110000092000125
3001 000111000111000102000105
3101 000068000108000108000081
3201 000048000096000121000044
3301 000042000078000123000020
3401 000034000073000125000017
3501 000048000084000125000014
-1
PUNCHED CORRELATION MATRIX OUTPUT
3333 4
3333 4
3333 4
3333 4
333323
333323
333323
333323
333321
60
1 1 0.1000000E
1 2 0.7225732E
1 3-0.5798441E
1 4 0.7999757E
1 1 0.1122857E
1 2 0.1030857E
1 3 0.1016857E
1 4 0.9960000E
1 1 0.3500000E
01 0.7225732E
00 0.1000000E
00-0.6567597E
00 0.8950214E
03 0.5141101E
03 0.2161068E
03 0.2601331E
02 0.7545321E
02
00-0.5798441E
01-0.6567597E
00 0.1000000E
00-0.7525222E
02
02
02
02
00 0.7999757E
00 0.8950214E
01-0.7525222E
00 0.1000000E
00
00
00
01
PUNCHED FACTOR SCORES OUTPUT
125 0.1952047E 01-0.2331642E 01
225 0.2963847E 00-0.1594353E 01
325-0.4612540E 00-0.1131953E 01
425-0.2389709E 01-0.7728281E 00
525-0.3506941E 01-0.5253528E 00
625-0.1578879E 01-0.5796633E 00
725-0.3429788E 00-0.5622065E 00
825 0.6738971E-02-0.5635983E 00
925 0.8490992E 00-0.5561192E 00
1025 0.3999518E-01-0.6032414E 00
1125-0.1011253E 01-0.5692308E-01
1225-0.2248286E 01 0.1632854E 00
1325-0.2765937E 01 0.1167946E 01
1425-0.1946555E 01 0.98P1128E 00
1525 0.1136240E 01 0.1907818E 00
1625 0.2493227E 01-0.4555233E 00
1725 0.1984453E 01-0.4183342E 00
1825 0.1837948E 01-0.2014065E 00
1925-0.1365285E 00 0.2584076E 00
2025-0.1992324E 01 0.5247329E 00
2125-0.3228456E 01 0.6974968E 00
2225-0.3100469E 01 0.2611022F 01
2325-0.3355663E 01 0.2021247E 01
2425-0.3044062E 01 0.1755776E 01
2525-0.3732885E 00 0.1551542E 01
2625 0.1751099E 01 0.10n5148E 01
2725 0.2006701E 01 0.8391422E 00
2825 0.2129229E 01 0.4882641E 00
2925 0.1194178E 01 0.3152475E Op
3025 0.4921422E 00-0.4477804E-01
3125 0.2693805E 01-0.3768844E 00
3225 0.3135445E 01-0.8926919E 00
3325 0.2451350E 01-0.9157380E 00
3425 0.2694848E 01-0.1001286E 01
3525 0.2337661E 01-0.9936804E 00
OUTPUT
61
// XEO FCTR
05
*LOCALFCTIR,FMTRNDATRDORNTBOXRAD,TRAN
*LOCALFCTR1.TRIDI.ORIONVRS
*LOCALCORELORNT
*LOCALFCTR20/ECTRORNT
*LOCALFCTR30/ARmX0R0mXIISCOPEOFOUT
FACTOR ANALYSIS SAMPLE PROBLEM
4
INPUT TYPE
SEQUENCE CHECK
VARIABLES ON CARD 1
VARIABLES ON CARD 2
VARIABLES ON CARD 3
TRANSFORMATION SWITCH
OUTPUT RAW CROSS PRODUCTS
0
0
0
0
0
OUTPUT
OUTPUT
OUTPUT
FACTOR
HUN KER
NUM BER
RESIDUAL CROSS PRODUCTS
VARIANCE
COVARIANCE
CORRELATION
SCORES
OF FACTORS OPTION
OF FACTORS OR PERCENT OF TRACE
1
2
2
2
2
COMMUNALITY OPTION
ROTATION OPTION
NUMBER OF FACTORS TO ROTATE
POOLING OPTION
LATENT VECTORS
UNROTATE0 FACTOR MATRIX
ORTHOGONAL TRANSFORMATION MATRIX
0
2
0
0
1
ORTHOGONAL FACTOR MATRIX
TRANSFORMATION MATRIX TO OBLIQUE REFERENCE VECTOR STRUCTURE
1
OBLIQUE REFERENCE VECTOR STRUCTURE MATRIX
CORRELATIONS AMONG OBLIQUE REFERENCE VECTORS
OBLIQUE REFERENCE VECTOR PATTERN MATRIX
1
1
1
CORRELATIONS BETWEEN REFERENCE VECTORS AND PRIMARY FACTORS
1
OBLIQUE PRIMARY FAC T OR STRUCTURE MATRIX
CORRELATIONS AMONG JBLIOUE PRIMARY FACTORS
OBLIQUE PRIMARY FACT^R PATTERN MATRIX
FACTOR SCORE. REGRESSION COEFFICIENTS
1
1
1
(212,1X.4F6.0/
62
JOB
NUMBER OF VARIABLES
1
1
1
3333
PAGE
0
FACTOR ANALYSIS SAM PLE PROBLEM
JOB
3333
PAGE
2
JOB
3333
PAGE
3
JOB
3333
PAGE
MATRIX OF RESIDUAL CROSS-PRODUCTS
VARIABLE
PI
02
P3
P4
04
P3
P2
01
0.89865E 05 0.27295E 05 m 0.26365E 05 0.10550E 06
0.77295F 05 0.15878E 05 m 0.12553E 05 0.49620F 05
-0.26365E 05 m0.12553E 05 0.23007E 05 -0.50219E 05
0.10550E 06 0.49620E 05 m 0.50219E 05 0.19356E 06
FACTOR ANALYSIS SAMPLE PROBLEM
VARIANCE - COVARIANCE MATRIX
VARIARLF
P1
02
03
04
P4
PI
P3
P2
0.26430E 04 0.80279E 03 m 0.77546E 03 0.31032F 04
0.40270E 03 0.46702E 03 -0.36920E 03 0.14594E 04
m0.77546E 03 m0.36920E 03 0.67669E 03 -0.14770E 04
0.31032F 04 0.14594F 04 -0.14770E 04 0.56931E 04
FACTOR ANALYSIS SAMPLE PROBLEM
SUMMARY STATISTICS
VARIABLE
1
2
3
4
P1
P2
P3
P4
NO.OF CASES •35
LOW
0.34000E
0.62000E
0.39000E
o.14000E-
02
02
02
02
AVERAGE
HIGH
0,23700E
0.15200E
0.15900E
0.36600E
03
03
03
03
0.11228E
0.10300E
0.101013E
0.99600E
03
03
03
02
STO. DEV.
0.51411E
0.21610E
0.26013E
0.75453E
02
02
02
02
VARIANCE
0.26430E
0.46702E
0.67669E
0.56931E
04
03
03
04
63
3333
PAGE
6
JOB
3333
PAGE
7
JOB
3333
PAGE
8
JOB
3333
PAGE
FACTOR ANALYSIS SAMPLE PROBLE MJOB
P1
P2
P3
P4
MATRIX OF CHARACTERISTIC VECTORS
-0.49911E 00 -0.54530E nn
..0.51266E 00 -0.16668E 00
0.46188E 00 -0.81965E 00
-0.53892E 00 -0.55082E-01
FACTOR ANALYSIS SAMPLE PROBLEM
4.0000
TRACE
CHARACTERISTIC ROOTS
3.2134
0.4301
0.2753
0.0810
CUMUL. PERCENT OF TRACE
80.3374
91.0900
0.0000
0.0000
FACTOR ANALYSIS SAMPLE PROBLEM
P1
P2
P3
P4
-0.96604E
-0.91901E
0.92799E
-0.96608E
on -0.35762E on
00 -0.10931E 00
On -0.53754E 00
00 -0.36124E-01
NORMALIZED UNROTATED FACTOR LOADINGS
COM M UNAL IT IFS
0.8779364E 00
0.8565298E 00
0.9745132E 00
0.9346225E 00
FACTOR ANALYSIS SAMPLE PROBLE M9
NORMAL VARIMAX CRITERION (NORMALIZED)
CYCLE
1
2
3
CRITERION
DIFFERENCE
0.02950188
0.18610206
0.18610206
0.02950188
0.15760016
0.00000000
EPSILON CRITERION.
JOB
FACTOR ANALYSIS SA M PLE PROBLEM
ORTHOGONAL TRANSFORMATION MAT•IX..
VARIABLE
1
7
64
1
0.7966
0.6044
2
-0.6044
0.7966
0.00116000
3333
PAGE
10
FACTOR ANALYSIS SAMPLE PROBLEM
JOB
3333
PAGE
11
JOB
3333
PAGE
12
3333
PAGE
13
JOB
3333
PAGE
14
JOB
3333
PAGE
15
ORTHOGONAL FACTOR MATRIXIVARIMAX)
VARIABLE
P1
P2
P3
P4
1
2
-0.9061
-0.7982
0.3346
-0.7914
0.2385
0.4684
-0.9287
0.5551
FACTOR ANALYSIS SAMPLE PROBLEM
TRANSFORMATION TO OBLIQUE REFERENCE VECTOR STRCTR.
VARIABLE
1
2
1
0.9322
0.3617
2
0.3853
0.9227
FACTOR ANALYSIS SAMPLE °ROBLE°
JOB
CORRELATIONS AMONG OBLIQUE REFERENCE VECTORS
VARIABLE
1
2
1
1.0000
0.6931
2
0.6931
1.0000
FACTOR ANALYSIS SAMPLE PROBLEM
OBLIQUE REFERENCE VECTOR STRUCTURE MATRIX
1
2
P1
-0.7584
-0.1290
P2
P3
P4
-0.5746
-0.0239
-0.5370
0.1246
-0.7279
0.2072
VARIABLE
FACTOR ANALYSIS SAMPLE PROBLEM
OBLIQUE REFERENCE VECTOR PATTERN MATRIX
VARIABLE
P1
P2
P3
P4
1
-1.2874
-1.2722
0.9249
-1.3099
2
0.7632
1.0063
-1.3690
1.1151
65
3333
PAGE
16
JOB
3333
PAGE
17
JOB
3333
PAGE
18
JOB
3333
PAGE
19
JOB
3333
PAGE
20
JOB
FACTOR ANALYSIS SAMPLE PROBLEM
CORR. BET. REFERENCE VECTORS AND PRIMARY FACTORS
VARIABLE
1
2
1
0.7208
0.0000
2
0.0000
0.7208
FACTOR ANALYSIS SAMPLE PROBLEM
CORR. AMONG OBLIQUE PRIMARY FACTORS
VARIABLE
1
2
1
1.0000
-0.6931
2
-0.6931
1.0000
FACTOR ANALYSIS SAMPLE PROBLEM
VARIABLE
P1
OBLIQUE PRIMARY FACTOR STRUCTURE MATRIX
- 0.9280
2
0.5502
1
P2
-0.9170
0.7254
P3
0.6667
-0.9868
P4
-0.9442
0.8038
FACTOR ANALYSIS SAMPLE PROBLEM
OBLIQUE PRIMARY FACTOR LOADINGS
VARIABLE
P1
P2
P3
P4
1
-1.0521
0.7972
-0.0332
.m0.7450
2
-0.1790
0.1728
-1.0098
0.2875
FACTOR ANALYSIS SAMPLE PROBLEM
FACTOR SCORE REGRESSION COEFFICIENTS
VARIABLE
P1
P2
PS
P4
66
1
-2.9865
0.9614
0.4487
0.8373
2
0.1404
-0.0652
- 1.0582
-0.0641
JOB
FACTOR ANALYSIS SAMPLE PROBLEM
3333
PAGE
21
FACTOR SCORES
IDENTIFICATION
1
1
2
2
3
3
4
4
5
5
6
6
7
7
8
8
9
9
10
10
11
11
12
12
13
13
14
14
15
15
16
16
17
17
18
18
19
19
20
20
21
21
22
22
23
23
24
74
25
25
26
26
27
27
28
28
29
29
30
30
31
31
32
32
33
33
34
34
35
35
1
0.19520E 01
0.29638E 00
0.461258 00
43.238978 01
0.350698 01
-0.15788E 01
■00342978 00
0.67389802
0.849098 00
0.399958..01
0.101128 01
13.224828 01
■0.276598 01
..0.19465E 01
0.11362E 01
0.24932E 01
0.19844E 01
0.18379E 01
..0.13652E 00
■0.199238 01
0.322848 01.
-.0.31004E 01
0.335568 01
0.304408 01
0.373288 00
0.17510E 01
0.20067E 01
0.21292E 01
0.11941E 01
0.49214E 00
0.269388 01
0.31354E 01
0.24513E 01
0.26948E 01
0.233768 01
2
■ 0.233168 01
■ 0.15943E 01
■ 0.113198 01
...0.772828 00
-0.52535E 00
..0.57966E 00
0.56220E 00
■0.56359E 00
13.556118 00
■0.603248 00
-0.56823E-01
0.1632SE 00
0.11678E 01
0.98811E 00
0.190788 00
0.45552E 00
..0.41833E 00
■0.20140E 00
0.25840E 00
0.52473E 00
0.69749E 00
0.26110E 01
0.202128 01
0.17557E 01
0.15515E 01
0.10051E 01
0.83914E 00
0.48826E 00
0.31524E 00
13.447788.-01
0.376888 00
■ 0.892698 00
■ 0.915708 00
.. 0.10012E 01
0.99368E 00
JOR COMPLETED
CORRELATION MATRIX INPUT--MULTIPLE R**2 ON DIAGONAL
// XED FCTR
05
*LOCALFCTRIFMTRD,DATRD,PRNTB.MXRAD,TRAN
*LOCALFCTR1,TRIDI.OR,INVRS
*LOCALCOREL,PRNT
*LOCALFCTR2,VECTR,PRNT
*LOCALFCTR3,VARMX,PROMX,SCORE,RFOUT
020200
3333
FACTOR ANALYSIS SAMPLE PROBLEM
040300000000000000000100320200020000 1010001010101000101010100
P2
P4
P1
P3
3333 4 1 1 0.6412571E 00 0.7225732E 00-0.5798441E 00 0.7999757E
3333 4 1 2 0.7225732E 00 0.8018028E 00-0.6567597E 00 0.8950214E
3333 4 1 3-0.5798441E 00-0.6567597E 00 0.5689992E 00-0.7525222E
3333 4 1 4 0.7999757E 00 0.8950214E 00-0.7525222E 00 0.8816457E
333323 1 1 0.1122857E 03 0.5141101E 02
333323 1 2 0.1030857E 03 0.2161068E 02
333323 1 3 0.1016857E 03 0.2601331E 02
333323 1 4 0.9960000E 02 0.7545321E 02
333321 1 1 0.3500000E 02
-1
00
00
00
00
OUTPUT
67
// XEQ FCTR
05
*LOCALFCTR,FMTRDOATRDORNTAOIXRANTRAN
*LOCALFCTR10TRINOR,INVRS
*LOCALCOREL,PRNT'
*LOCALFCTR2oVECTR$PRNT
*LOCALFCTR30/ARMX,PROMX.SCORFoRFOUT
FACTOR ANALYSIS SAMPLE PROBLEM
JOB
3333
JOB
3333
PAGE
1
JOB
3333
PAGE
2
PAGE
0
NU M BER OF VARIABLES
IN°UT TYPE
SEQUENCE CHECK
VARIABLES ON CARD 1
VARIABLES ON CARD 2
VARIABLES ON CARD 3
TRANSFORMATION SWITCH
OUTPUT RAW CROSS PRODUCTS
OUTPUT RESIDUAL CROSS PRODUCTS
OUT P UT VARIANCE
COVARIANCE
OUTPUT CORRELATION
F ACTOR SCORES
NU M BER OF FACTORS OPTION
NUMBER OF FACTORS OR PERCENT OF TRACE
COMMUNALITY OPTION
ROTATION OPTION
NUMBER OF FACTORS TO ROTATE
0 00LING OPTION
LATENT VECTORS
UNROTATED FACTOR MATRIX
ORTHOGONAL TRANSFORMATION MATRIX
ORTHOGONAL FACTOR MATRIX
TRANSFORMATION MATRIX TO OBLIQUE REFERENCE VECTOR STRUCTURE
OBLIQUE REFERENCE VECTOR STRUCTURE MATRIX
CORRELATIONS AMONG OBLIQUE REFERENCE VECTORS
OBLIQUE REFERENCE VECTOR PATTERN MATRIX
CORRELATIONS BETWEEN REFERENCE VECTORS AND PRIMARY FACTORS
OBLIQUE PRIMARY FACTOR STRUCTURE MATRIX
CORRELATIONS AMONG OBLIQUE PRIMARY FACTORS
OBLIQUE PRIMARY FACTOR PATTERN MATRIX
FACTOR SCORE REGRESSION COEFFICIENTS
FACTOR ANALYSIS SAMPLE PROBLEM
MATRIX OF CHARACTERISTIC VECTORS
P1
P2
P3
P4
■0.46728E
-0.52362E
0.43527E
-0.56391E
00 .0.47333E 00
00 ..0.31961E 00
00 II).81887E 00
00 0.56930E..01
FACTOR ANALYSIS SA T. LE ORONLEM
TRACE
2.0937
CHARACTERISTIC ROOTS
68
CUMUL. PERCENT OF TRACE
2.9564
102.1689
0.0298
..0.0061
-0.0864
103.1991
0.0000
0.0000
FACTOR ANALYSIS SAMPLE PROBLEM
JOB
3333
PAGE
3
JOB
3333
PAGE
4
JOB
3333
PAGE
5
JOB
3333
PAGE
6
PAGE
7
NORMALIZED UNROTATED FACTOR LOADINGS
-0.80146E
-0.00033E
0.74842E
-0.96961E
P1
02
P3
P4
00 -0.81725E-01
00 -0.55183E-01
00 -0.14118E 00
00 0.98294E-02
COmmUNALITIFS
0.6522407F 00
0.8136515E 00
0.5801334E 00
0.9407521E 00
FACTOR ANALYSIS SAMPLE PROBLEM
NORMAL VARIMAX CRITERION (NORMALIZED/
CRITERION
0.00035874
DIFFERENCE
0.00035874
2
0.02373140
0.02337265
3
0.02373140
0.00000000
CYCLE
1
EPSILON CRITERION.
FACTOR ANALYSIS SAMPLE PROBLEM
0.00116000
ORTHOGONAL FACTOR MATRIXIVARIMAXI
-0.6504
2
0.4786
P2
0.7044
0.5633
P3
P4
0.4599
-0.7121
0.6071
0.6580
VARIABLE
P1
1
FACTOR ANALYSIS SAMPLE PROBLEM
TRANSFORMATION TO OBLIQUE REFERENCE VECTOR STRCTR.
VARIABLE
1
0.8730
0.4876
1
2
2
0.3683
0.9296
FACTOR ANALYSIS SAMPLE PROBLEM
JOB
3333
CORRELATIONS AMONG OBLIQUE REFERENCE VECTORS
VARIABLE
1
2
1
1.0000
0.7749
2
0.7749
1.0000
69
FACTOR ANALYSIS SAMPLE PROBLEM
JOB
3333
PAGE
8
JOB
3333
PAGE
9
JOB
3333
PAGE
10
JOB
3333
PAGE
11
JOB
3333
PAGE
12
OBLIQUE REFERENCE VECTOR STRUCTURE MATRIX
VARIABLE
P1
02
03
04
-0.3344
-0.3403
2
0.2054
0.2642
0.1054
-0.3008
-0.3950
0.3494
1
FACTOR ANALYSIS SAMPLE PROBLEM
CORR. BET. REFERENCE VECTORS AND PRIMARY FACTORS
VARIABLE
1
0.6320
0.0000
1
2
2
0.0000
0.6320
FACTOR ANALYSIS SAMPLE PROBLEM
CORR. AMONG OBLIQUE PRIMARY FACTORS
VARIABLE
1
2
1
1.0000
-0.7749
2
-0.7749
loonn
FACTOR ANALYSIS SAMPLE PROBLEM
OBLIQUE PRIMARY FACTOR STRUCTURE MATRIX
VARIABLE
1
2
01
-0.7810
0.7351
P2
-0.8624
0.8353
P3
P4
0.6512
-0.9045
(:).7543
0.9218
FACTOR ANALYSIS SAMPLE PROBLEM
OBLIQUE PRIMARY FACTOR LOADINGS
VAR/ABLE
01
P2
-0.5291
2
0.3250
0.5384
0.4181
P3
0.1668
-0.6250
P4
-0.4760
0.5529
JOB COMPLETED
70
1
2. 2. 9 References
Businger, P.A. "Eigenvalues of a real symmetric matrix by the QR method",
Algorithm 253, Communications of the Association for Computing Machinery,
April 1965, vol. 8, no. 4.
Guttman, L. "Some Necessary Conditions for Common Factor Analysis",
Psychometrika, 1954, 19, 149-162.
Harman, H. Modern Factor Analysis. University of Chicago Press, 1960, pp. 289-308;
pp. 261-288; pp. 349-356.
Hendrickson and White. "Promax: A quick method for rotation to oblique simple
structure", British Journal of Statistical Psychology, 1964, XVII, 65-70.
Horst, P. Factor Analysis of Data Matrices. Holt, Rinehart, and Winston, Inc.
New York, Chicago, San Francisco, Toronto, London. 1965, p. 214.
HOW. FORTRAN Subroutine, using non-iterative methods of Householder, Ortega,
and Wilkinson, solves for eigenvalues and corresponding eigenvectors of a real
symmetric matrix. Program Writeup F2 BC HOW. Berkeley Division, University of
California, 1962.
Kaiser, "The Varimax criterion for analytical rotation in factor analysis",
Psychometrika, 1958, 23, 187-200.
Wilkinson, J. H. "The calculation of the eigenvectors of codiagonal matrices", The
Computer Journal, 1958, vol. 1, p. 90.
Wilkinson, J.H. "Householder's method for the solution of the algebraic eigen
problem", The Computer Journal, 1960, vol. 3, p. 23.
71
2.3 ANALYSIS OF VARIANCE
From the experimental observations on a variable x, this program will compute an
analysis of variance for a complete factorial design for a maximum of four (4) factors.
The method used in the program is essentially that described by H.O. Hartley. * This
method is particularly useful, since it can be extended to accommodate a great many
experimental designs.
The extension to other experimental designs is accomplished by a very simple
procedure. The program performs a factorial analysis and then allows the user to pool
certain components of the analysis of variance table in accordance with the summary
instructions that specifically apply to the particular design desired. For example, a
two- or three-factor design can result in the following analysis of variance tables:
•
Single classification
•
Two-way classification with cell repetition
•
Randomized block with two factor treatments
•
Split plot
•
Split-split plot
•
Three-factor randomized blocks
By utilizing a special report generator, the user has flexibility in choosing the
appropriate components to pool in forming the error term or terms to accommodate
the above designs or any other similar designs.
Once the data is contained in storage, the sum of squares is computed as follows:
A. = Sum of all the observations at level a.
1
Let
n = Number of observations summed to obtain A.
a
1
AB.. = Sum of all observations at level ab..
1J
LJ
nab
= Number of observations summed to obtain AB. .
ABC
n
ijk
abc
= Sum of all observations at level abc
ijk
= Number of observations summed to obtain ABC
ijk
Thus, a general formula for the main effect due to factor A is:
A2
SS =
a
i
na
-
G
n
2
*Ralston, A. and Wilf, H. S., Mathematical Methods for Digital Computers. New York:
John Wiley and Sons, Inc. , 1960, pp. 221-230.
72
where G = the grand total of all observations, and li e is the number of observations
summed to obtain G. The main effect due to factor B has the form:
SS
b
2
ZB
i
= nb
G
n
2
g
The general computational formula for the variation due to the AB interaction is:
Z (AB i .)
SS
ab
=
n
2
G
2
(SS + SS )
a
b
ab
The general computational formula for the variation due to the ABC interaction is:
Z (ABC .., )
SS
abc
-
2
11 K.
n
G
2
abc
(SS + SS + SS + SS
+ SS
+ SS )
a
b
c
ab
ac
bc
The following table shows the layout of the complete table for the analysis of variance
of a complete factorial design:
ANOVA TABLE FOR A COMPLETE FACTORIAL DESIGN
Source of Viariation
A
Main effect
B
Main effect
Main effect
C
AB Interaction
AC Interaction
BC Interaction
ABC Interaction
Experimental error
(within cell)
Total
D. F .
(p-1)
(q-1)
(r-1)
(p-1)(q-1)
(p-1)(r-1)
(q-1)(r-1)
(p-1)(q-1)(r-1)
pqr (n-1)
npqr-1
Sum of
Squares
SSa
SSb
SS
SSab
Mean Square
SSa/ (p-1)
SSac
SSbe
SSabe
SSb/ (q-1)
SSe/ (r-1)
SSab/ (p-1)(q-1)
SSac/ (p-1)(r-1)
SSbe/ (q-1)(r-1)
SSabc/ (p-1)(q-1)(r-1)
SSerror
SSerror/ P qr (n-1)
SStotal
where p = number of levels in factor A
q = number of levels in factor B
r = number of levels in factor C
n = number of observations per cell
To obtain other experimental designs from a complete factorial design, the user should
analyze the data as if it were a complete factorial design, and then reconstruct his
ANOVA table from the output.
Not all experimental designs can be handled by this technique, notably Latin and Youden
squares, lattices, and incomplete randomized blocks.
73
Also, this program does not handle repeated measurement designs (that is, replications
must be considered as a factor). For a detailed account of the various experimental
designs, see 0. Kempthorne, Design and Analysis of Experiments (John Wiley, 1952).
Single Classification Design (A X B). In this case, the replications are considered as a
factor (B). The error term is:
SSerror = SS b + SSab
giving the following reconstructed ANOVA table:
SSa
SSerror
SStotal
Two-Way Classification with Cell Repetition (A X B X C). This differs from a randomized block design in that one is not interested in the recovery of interblock information.
Consequently, the error term is:
SS= SS+ SS + SS + SS
error
abc
c
ac
bc
(where factor C is the cell repetitions) thus giving the following ANOVA table:
SSa
SSb
SSab
SSerror
SStotal
Randomized Block with Two Treatment Factors (A X B X C). Here the third "factor"
(C) is blocks. In this case, one is interested in finding out whether there are significant
differences between blocks, so the error term is computed from:
SS
error
= SS
ac
+ SS
bc
thus giving an ANOVA table as follows:
SSa
SSb
SSab
SS (blocks)
SSerror
SS total
74
+ SS
abc
Split-Plot Design (A X B X C). In this case, let factor A = main treatments; B = subtreatments, and C = blocks. Then, appropriate error terms are calculated as follows:
(a) SSerror =SS ac
(b) SSerror = SSbc+ SSabc
The ANOVA table becomes:
Main treatment
A
SS
Blocks
C
SS
Error (a)
SS
Subtreatment
Interaction
B
SS
AXB
SS
Error
SS
Total
SS
a
c
ac
b
ab
bc
+ SS
abc
total
Split-Split Plot Design (A X B X C X D). The factors in this case are:
A = Main treatment
B = Subtreatment
C = Sub-subtreatment
D = Blocks
Consequently, there will be three separate error terms:
(a) SS
error
= SS
ac
(b) SSerror = SSbc+ SS abc
(c) SS
error
= SS
dc
+ SS
dcb
+ SS
dca
+ SS
dcab
This gives the following reconstructed ANOVA table:
A
SS
C
SS
B
SS
a
c
Error (a) SSac
b
75
A X B SS
ab
Error (b) SSbc +
D
SS
SSabc
d
A X D SS
ad
B X D SSbd
AXBXD SS
abd
Error (c) SS _ c +
d SSdcb SSdca + SSdcab
Total
SS
total
Three-Factor Randomized Blocks (A X B X C X D). Let factor C = blocks.
The error term becomes:
SS
error
= SS + SS + SS + SS + SS + SS + SS
bcd
abcd
bc
dc
abc
acd
ac
Thus giving the following reconstructed ANOVA table:
A
SS
B
SS
a
b
SS
d
SS
C (blocks)
c
SS
AXB
ab
SS
AXD
ad
SS
BXD
bd
AXBXD SS
abd
+ SS
+ SS
+ SS
SSac + 5519c + SS + SS
Error
abed
abc
bed
acd
dc
D
Total
SS
total
2.3.1 Tests of Significance
The output of this program consists of the sums of squares and mean squares for all
the main effects and interactions, together with the error mean square. In general,
these main effects and interactions are tested for significance by dividing the mean
square for the particular effect or interaction by the appropriate error term. The
difficulty arises in the choice of the appropriate error term. A brief account of how
to choose the correct error term is given below. This account is by no means comprehensive, and if the user is in any doubt as to the error term to use in his own case,
he should consult H. Scheffe, The Analysis of Variance, John Wiley, 1959 .
76
The basis for the choice of error term in ANOVA F-tests is the type of structural model
used for the analysis of variance. Three models are discussed below; these should
cover the majority of cases.
1. Model I (Fixed Effects). The fixed effects model is applicable when the factors used
in the experiment include all possible levels for each factor, and when inferences
are not made about any levels not included. Examples of this are such factors as
sex, where there are only two levels possible; or training methods for teaching a
specific skill; or, in a drug experiment, the treatments factor, where one is interested in the drugs used, and would not obviously want to make inferences to other
drugs not included in the experiment. In the case of a fixed factor, the investigator
is interested only in the levels of the variable studied in the experiment and not in
any others. In this case, the computation of the F-ratio is relatively simple. The
F-ratios are calculated using the error term as a divisor (for example, MS /MS
a
error
for the A main effect;
AB interaction, etc.).
MSab/MSerror:
2. Model II (Random Effects). The random effects model applies when the experiment
involves only a random sample of the set of treatments about which the experimenter
wants to make inferences. For example, to study the effects of a certain drug (say
alcohol) on driving skill, one would have several different levels (doses) of alcohol
within the drug factor. However, all possible levels of alcohol could not be used, so
one takes what is considered to be a random sample of the levels within the factor
and then makes inferences about other levels.
Another example would be the following:
To study the effects of level of illumination on productivity in a factory, the luminance
factor would be a random effects factor, since all possible levels of luminance would
not be used in the experiment, but only a sample of them.
An analysis of variance with all random effects is rarely found, and the calculation
of F-ratios for this case presents some difficulties. For a two-factor model (A X B),
the F-ratios are:
A: MSa/MSab
B: MSb/MSab
AB: MS
ab
/MS
error
For the three-factor case (A X B X C), we have the following F-ratios:
A : (MSa
MSabc)/(MSac
MSab)
B: (MS + MS
)/(MS + MS )
abc
b
ab
bc
C: (MSc
MSabc)/(MSac
MSbc)
77
AB: MS
ab
AC : MS
ac
BC: MS
bc
/MS
/MS
/MS
abc
abc
abc
ABC: MS /MS
/MS
error
In the case of the random effects model, the interactions should be tested for
significance first, because if they are found to be significant, there is little point
in testing the main effects for significance.
3. Model HI (Mixed Model). The mixed model is probably the most common form used
in analysis of variance. Here some factors are fixed, and others are random.
The calculation of the F-ratios in this case depends on which factors are fixed and
which are random. An example is given with two fixed factors and one random
factor.
A X B X C Design with Factor A a random effect:
*A : MS /MS
/MS
error
B : MS /MS
b
ab
C : MS /MS
c
ac
*AB: MS
ab
/MS
*AC: MS /MS
/MS
BC : MS
ABC: MS
bc
/MS
abc
error
error
abc
/MS
error
2. 3. 2 Job Execution
To perform an analysis of variance, the user must supply four sets of cards to the
program:
1. Monitor control cards
2. Program control cards
3. Data cards
4. Table output specification cards
*Random effects
78
Monitor Control Cards
The monitor control cards are necessary to initiate program loading from the disk and
to establish the necessary communication with the monitor. A general description of
the cards may be found in IBM 1130 Disk Monitor System Reference Manual (C 26-3750).
An analysis of variance requires the following cards:
14,
1 4 8
CC:
I
16-17
// XEQ ANOVA 02
*LOCALANOVA, FMTRD, PRNTB, DATRD, STORE
*LOCALANOV2, SDOP, MNSQ, REPRT
The monitor control cards do not change from job to job, but must be included with every
job processed.
Program Control Cards
The program control cards communicate the data-specific parameters and output options
to the program. The five card types are described below. In addition, control cards are
necessary for defining the format and content of the ANOVA table (section 2.3.3).
1. Input/output units card*
2. Job-title card*
3. Option card (described below)
4. Variable format card*
5. Table generation card (section 2.3.3)
Option Card
Number of Factors (cc 1-2)
This field is punched with an integer, n, less than or equal to 4; n is the total number
of factors in the experiment.
Input Mode (cc 3-4)
This field must be punched with an integer, n, which may take the values 1 or 2. If n is
equal to 1, the program reads the raw data from the 1442 card reader. If n = 2, the raw
data is read from the disk, where it has previously been transferred by a program using
input mode number one. The data is retained until destroyed by input (mode 1) from one
of the four system programs.
*See "General Operating Instructions", Chapter 1.
79
Transformation Switch (cc 5-6)
If the value in this field is equal to zero, the transformation program is not used. If
the value is 1, the transformation program is called after each data item has been read,
and before the item is stored in the appropriate storage cell. The transformation
program itself is a user-written FORTRAN program, which is discussed in section 2.5.1.
Number of Levels for Factor 1 (cc 7-8)
This field must be punched with an integer, n, which indicates the number of levels in
the first factor. For example, n would be equal to three for a 3X4X5 factorial design;
n should be less than ten.
Number of Levels for Factor 2 (cc 9-10)
Same as cc 7-8. This field is for factor 2.
Number of Levels for Factor 3 (cc 11-12)
Same as cc 7-8. This field is for factor 3.
Number of Levels for Factor 4 (cc 13-14)
Same as cc 7-8. This field is for factor 4.
Columns 11-12 and 13-14 may be left blank for a two-factorial experiment. However,
the program does not operate for fewer than two factors. All four factors, whether used
or not, must be accounted for on the variable format card. The product of the levels of
the factors is limited to 2000.
Analysis of Variance Option Card Summary
Column
Meaning
1-2
Number of factors
3-4
Input Mode
1 - Source data from card reader
*2 - Source data from disk
5-6
Transformation Switch
0 - No transformation
1 - Transformation
7-8
Number of levels for factor 1
9-10
Number of levels for factor 2
11-12
Number of levels for factor 3
13-14
Number of levels for factor 4
*Data previously entered under mode 1 is available for mode 2 until destroyed by input
(mode 1) from this or one other of the system's four main programs.
80
2.3.3 Analysis of Variance Table Generation
The operation of the program is designed to handle a general four-factorial design. The
program will read the data, form the deviates, and accumulate the sums of squares as
if all four factors were always present. As a result of this operation, certain accumulation and storage areas, in general, have cells that are not used unless all four factors
are present. In forming the analysis of variance table for a particular design, the user
has the option of pooling component sums of squares to form the error sums of squares
specific to the design. The sums of squares are located in storage and can be accessed
by the table generator cards. The table, components, and index are shown below:
Subscript
Component
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
A
AB
AC
AD
BC
BD
CD
ABC
ABD
ACD
BCD
ABCD
For a two-factor experiment, the sums of squares are located in cells with subscripts
1,2,5
For a three-factor experiment, the sums of squares are located in cells with subscripts
1,2,3,5,6,8,11
For a four-factor experiment, the sums of squares are located in cells with subscripts
1,2... ,15
Table Generator Card Format
To print the proper component and compute the appropriate error term for a particular
design, a set of table cards indicating the appropriate terms must be punched. A
description of this card is given below:
Column
1-16
Meaning
Row heading for component. This field may contain any 16 or
fewer characters that serve to identify the row of the table.
81
Column
Meaning
17-20
Print and end control.
0 or blank - Normal card.
+1 - Skip to a new page and print column headings and title
information before printing line for this component.
-1 - Last card. No more component cards will follow. The
analysis will be complete after this card is processed.
The residual and total line will be printed.
21-22 Subscript of the correspondence table component, to be printed
or used in accumulation of the sums of squares. For example,
if this field contains a 5, the component AB will be used for
either printing (the remainder of the card is blank) or accumulation.
23-24
Subscript of the cell in the correspondence table to be added to the
cell used in cc 21-22. This procedure is used for adding components
to form the sums of squares. For example, if cc 21-22 contained a
1 and cc 23-24 contained a 4, the printed sums of squares, mean
square, and degrees of freedom would be A + AB.
25-26
27-28
The remaining two-digit fields have the same effect as cc 23-24,
but are used to add additional components before printing the line.
For example, if cc 21-22 contained a 2, cc 23-24 contained a 6,
and cc 25-26 contained a 9, the sums of squares, mean square,
and degrees of freedom would be printed as the cumulative summary
of B + AC + BD.
49-50
Table Generator Card Summary
Column
1-16
Meaning
Alphameric heading for analysis of variance component.
17-20
END of table indicator
0 - More cards to follow
-1 - No more cards to follow
1 - Skip to a new page before line is printed
21-22
Table component to be printed
23-24
Table component to be pooled
N - Add the component with subscript N (corres. table) to the
first component defined in cc 21-22.
82
25-26
Same as cc 23-24
27-28
Same as cc 23-24
Column
Meaning
29-30
Same as cc 23-24
31-32
Same as cc 23-24
33-34
Same as cc 23-24
35-36
Same as cc 23-24
37-38
Same as cc 23-24
39-40
Same as cc 23-24
41-42
Same as cc 23-24
43-44
Same as cc 23-24
45-46
Same as cc 23-24
47-48
Same as cc 23-24
49-50
Same as cc 23-24
2.3.4 Data Input
To set up the data for the analysis of variance, the user must identify each item of data
as to its factor and level, and punch this information on a card along with the data item.
Hence, each data card will have five fields, as follows:
Type
Meaning
1
Integer (I)
Number of level - factor 1
2
Integer (I)
Number of level - factor 2
3
Integer (I)
Number of level - factor 3
4
Integer (I)
Number of level - factor 4
5
Floating point (F)
Observation
Field
Disk working storage allows input of 499 observations.
The particular columns occupied by each field are arbitrary. The user describes the
format of the card by means of a variable format card, which is entered into the program
behind the option card. On the format card, provision must be made for all four "level"
fields, even though all four fields are not necessary in the particular analysis. Figure 11
shows a sample data card from a two-factor design.
83
1 2 3 4 5 6 7
CC:
Data card
2
4
11-12
32
1
CC:
Format card
(212, 211, F6.0)
I
Figure 11. Sample cards from two-factorial design
In normal usage, the data items are punched one to a card, with the appropriate
identification. Following the data deck, there must be an end-of-deck indicator card,
which is a card containing a negative number in the first field. The order of cards is
arbitrary, as the cards are rearranged in proper order before the analysis takes place.
2.3.5 Operating Instructions
A. Using the analysis of variance program when the total 1130 Statistical System has
not been stored on the disk
If the user wishes to load only the set of programs that allow analyses of variance,
the following programs must be compiled or assembled and stored on the disk. Each
deck begins with a card punched as
//FOR
and ends with an
*STORE
card.
The user should use a disk containing the 1130 Disk Monitor System, as described
in section 1.1. The following decks should be preceded by a cold start card, placed
in the card reader hopper, and the buttons IMMEDIATE STOP (console), RESET
(console), START (card reader), and PROGRAM LOAD (console) should be pressed.
A blank card should be placed after the last deck in the card reader hopper.
DECKS-LABELS: ANOVA-NOVA; STORE-STOR: GET-GETO; ANOV2-NOV2;
SDOP-SDOP; MNSQ-MNSQ; REPRT-RPRT; *FMAT-FMAT; *FMTRD-FMRD;
*DATRD-DTRD; *GMPYX-GMPY; *GDIVX-GDIV; *PRNTB-PRNB; TRAN-TRAN.
*Used in all four analysis types
84
B. Execution from Disk
Once the component subroutines and main calling programs are on the disk, the
execution of a job requires the monitor control cards, program control cards, and
data cards to be placed in the card reader. The deck should be preceded by a cold
start card. To initiate processing, the buttons IMMEDIATE STOP and RESET
(console), START (card reader), and PROGRAM LOAD (console) should be pressed.
The order in which the cards are placed in the card reader for either matrix or raw
data input is shown in Figures 12 and 13.
(
End of Data
Table Generation Deck,
-*including the end-of-deck
indicator*
Data Deck
Variable
(Format
Option
Job-Title
Input/Output
(Units
Monitor
Control
*Final table generator card
Figure 12. ANOVA — card reader input
85
Table Generation Deck,
.E.-including the end-of-deck
indicator*
(Option
Job-Title
/Input/Output
Units
Monitor
Control
*Final table generator card
Figure 13. ANOVA — disk input
2.3.6 Sample Problem
The data for this sample problem was taken, with permission, from page 276,
Statistical Theory in Research, by R. L. Anderson and T. A. Bancroft.
McGraw-Hill Book Company, Inc. , New York, 1952.
INPUT
// XEQ ANOVA
02
*LOCALANOVA,FMTRD,PRNTB,DATRD,STORE
*LOCALANOV2, SDOP,MNSQ,REPRT
020200
3333 TEST ANOVA-I
030101060302
(411,F4.0)
111
161
112
192
121
145
122
232
131
172
132
227
211
166
212
253
221
231
222
231
204
231
232
214
311
113
312
208
321
131
322
190
331
104
332
144
411
103
412
171
158
421
86
422
431
432
511
512
521
522
531
532
611
612
621
622
631
632
171
135
146
132
196
176
242
178
186
180
198
216
238
175
230
BLOCKS
FERTILIZER
VARIETY
F X V
ERROR
0101
02
03
08
—1050611
OUTPUT
// XEO ANOVA
02
*LOCALANOVA,FMTRDORNTBOATRDoSTORE
*LOCALANOV2,SDOP.MNSOoREPRT
ST ANOVA...!
NUMBER OF FACTORS
INPUT MODE
TRANSFORMATION SWITCH
NUMBER OF LEVELS — FACTOR
NUMBER OF LEVELS •• FACTOR
NUMBER OF LEVELS • FACTOR
NUMBER OF LEVELS — FACTOR
1
2
3
4
3
1
1
6
3
2
0
(4I1,F4.0)
87
ANALYSIS OF VARIANCE TABLE FOR
COMPONENT
SUM OF
SQUARES
3 X
2 X
DEGREES OF
FREEDOM
0 EXPERIMENT
MEAN
SQUARE
BLOCKS
FERTILIZER
24938.91
4034.00
5
2
4987.78
2017.00
VARIETY
17292.25
1
17292.25
1442.66
9896.91
2
25
721.33
395.87
57604.74
35
F X V
ERROR
TOTAL
JOB COMPLETED
88
6 X
2.4 LEAST-SQUARES CURVE FITTING BY ORTHOGONAL POLYNOMIALS
Given m points (Xi , Y i ), i= 1,
, m, the (X) set not necessarily being equally spaced,
this program will determine a polynomial of specified degree n or less,
2
y=a + a x + a x +
o
1
2
+ a
n
nx
which best approximates these points in the least-squares sense; n should not be
specified greater than ten; m, of course, must be greater than n, and for practical
purposes should be considerably greater than n. The program allows the number of
points, m, to be as high as 149. It is difficult to envisage a requirement such that
n = 7 will not suffice; however, the program has been successfully tested on polynomials
of order 16 with 134 data points, witii accurate results.
To maintain maximum accuracy, the program uses orthogonal polynomials, as described
by G.E. Forsythe. * The process of finding the polynomial is accomplished by beginning
with a first- and a second-degree polynomial and evaluating a variance criterion to
determine whether the second-degree will offer a better fit than the first. If the
variance criterion is satisfied within a specified tolerance, the program accepts the
second-degree polynomial computed to be the best fitting polynomial. If not, the nextorder polynomial is computed and compared to the second-degree. The process continues until the variance criterion is satisfied or a specified maximum degree reached.
The degree of the last polynomial is assumed to be the best fitting polynomial for the
data.
The essential characteristics of the method are as follows:
Let
y =
E
j=0
c.P. (x)
JJ
(1)
where each P. (x) is a polynomial of degree j.
By minimizing
2
M
i=1
[Yi
j =
0
cP (x.)]
jj
(2)
and letting
r.
jk
= E
1=1
S.
=
P (x) P (x.)
j
i ki
E
y.P. (x.)
(3)
(4)
i=1
*Forsythe, G.E., "Generation and Use of Orthogonal Polynomials for Data Fitting
with a Digital Computer", J. Soc. Indust. Appl. Math. 5, 1957; 74-78.
89
a set of normal equations are obtained
S. = E c r. ; j = 0,1,2,...n.
kk
k= 0
(5)
Equation (5) consists of a set of n + 1 simultaneous equations with n + 1 unknowns.
However, as the polynomials P. x) are orthogonal, then,
3(
0
for j k
rjk
=M
E
i=1
(6)
2
P. (x.) for j = k
Under this condition, the coefficients c. can now be evaluated by
S.
c. =
3 T..
33
j = 0,1,2,...,n
(7)
The polynomials P. x) are defined recursively by
3(
P (x) = 0
-1
Po(x)
P (x)
1
=1
= (x - a 1) P o(x) -Q 013_1(x)
(8)
•
Pj+1 (
= (x - 3+1
a . ) P.(x) - P. (x)
3 -1
where
a j+1
13j
90
i=1
2
x. P. (x.)
1 1
Li P.2(x.)
J
1
i=1
j = 0, ...n
j = 1, 2, ...n - 1 (00=0)
(9)
(10)
The polynomial solution vectors J' a3, and ^ are then used to compute a., the
j
coefficients of the equation
2
n
y = a o + a x + a 2 x + ... a
l
nx
for the degree n.
In principle, the coefficients a i could be used in equation (11) to compute the fitted
values for any given argument array. However, a il may change rapidly as n changes;
therefore, it would be necessary to compute a i with great precision. To avoid this
from equation (8), and
difficulty, fitted values use ci , a • and . to compute P•(x)
simultaneously compute the fitted values from equationJ (1).
Similarly, the derivative computation usesJ'c a . and . to compute the Al derivative
133
of. from the recurrence relation
PJ(x)
r
d rP
d P j +1
d r-1 Pd r P
j -1
- (x-a. )
(12)
.
j+1
r + r r-1 cbc r
J
r
dx
dx
dx
r= 1,2 ..... n
j = 0,1,2, ...n-1
for any given set of arguments.
If n and/or the range of x is large, the elements of the orthogonal vectors generated
change rapidly in size; this imposes severe restrictions on accuracy. This becomes
evident to the user by viewing the changes in the elements of successive vectors Pj(X)
as j increases, or by viewing residuals, which in this case may tend to increase
rather than decrease as j increases.
To aid in circumventing this problem, the user is allowed to elect, on option, to have
the program transform X to X' such that X' is in the range (-2, 2). This transformation
will cause elements of PJP
. (") to remain approximately uniform in size as j increases.
The transformation used is
x' = [4x-2 (xa)
+
x(m) )] / (x(m) - xo.) ) ,
(13)
th
where x(i) is the .-order
statistic from the set (x).
When this scaling is used, the following points must be considered:
1. The values of ci and Pj(x') (equation (2)) are calculated and presented using x'. Since
the y's are not transformed, the y* T s (estimated y's) and residuals calculated at x'
are the same as those which would have been calculated at x if c. and P . had been
obtained without transforming.
91
2. The coefficients a, from
n
= E a .x
i=0
are actually a', from
y = E ai(x') 1
i=0
(14)
3. If, in addition to transforming x, the user elects to punch a, /3 , and c for later use
in obtaining y*'s from a new set of x's, that is, for later use in prediction, the
transformation is retained, and the new x's will be transformed as were the original
x's; y*'s for these new x's will be the same as the estimated y's would have been if
transformations had not been performed.
4. Derivatives are calculated using equation (14).
dx'
_
dx dx' dx
5. The elements ofPl(x
• ' will
) now, as j increases, be of approximately the same size
for all j. The elements of P i (x') are related to those of P 3 (x) by the factor:
J
4
[X(m) -
X(1)]
2.4.1 Summary of Output
if scaling is elected).
1.
(x, y) for all cases (x', y
2.
Predicted value of y for 3 rd , 7th,
kth_order polynomials.
:3. Residuals for 3, 7, ..., kth -order polynomials, that is, if k=13, Y-Ypred. is given
for k . 3, 7, 11, and 13.
4. Orthogonal polynomials of all degrees to k.
Polynomial solution vectors a ,
, and c — used to generate orthogonal polynomials.
6. Coefficients of fitted polynomial.
7. Predicted values for externally supplied data set.
8. kth -order derivatives of the polynomial at user-specified points, where k is the
order of the polynomial.
9. Scaling equation used, if required.
10. Analysis of Variance Table.
92
2.4.2 Job Execution
To perform a polynomial regression analysis, the user must supply three sets of cards
to the program:
1. Monitor control cards
2. Program control cards
3. Data cards
Monitor Control Cards
The monitor control cards are necessary to initiate program loading from the disk and
to establish the necessary communication with the monitor. A general description of
the cards may be found in IBM 1130 Disk Monitor System Reference Manual (C26-3750).
The orthogonal polynomial program requires the following cards:
CC:
14
8
4
16
/1 XEQ POLY 02
*LOCALPOLY, TRAN, DATRD, FMTRD, PRNTB
*LOCALPOL2, POLSQ, PCOEF , PDER , P FIT
Monitor control cards do not change from job to job, but must be included with every
job processed. The first program operated on by this system should be preceded by a
cold start card.
Program Control Cards
The program control cards communicate the data-specific parameters, and output options
to the program. The four possible card types are described below.
1. Input/output units card*
2. Job-title card*
3. Option cards (described below)
4. Variable format card*
TYPE I OPTION CARD
This type of option card is to be used when data is being entered into the program for
the initial computation of the best fitting polynomial.
*See "General Operating Instructions", section 1.2.
93
Maximum Degree of Polynomial (cc 1-2)
This field should be punched with an integer, n, which is less than or equal to ten. The
program attempts to fit a polynomial to the data points until the variance criterion
punched in columns 17-26 is satisfied by successive-degree polynomials. If the
variance criterion is not satisfied when the degree of the polynomial reaches n, the
program prints a message to this effect and continues, using the solution to the
nth-degree polynomial. The value of n should be less than m+1, where m is the
number of the data points read.
Input Source (cc 3-4)
This field must be punched with an integer, n, which assumes a value of either one (1)
or two (2).
If n is equal to 1, the data points, followed by a negative identification card, are read
from the 1442 card reader. If n is equal to 2, the data points are read from the disk,
having previously been transferred there by a program using input mode 1. The data
on the disk is destroyed by any program using input mode 1. As noted below (Secondary
Input Sources), input type 3 also destroys disk data.
Coefficients of Fitted Polynomial (cc 5-6)
In the determination of the best fitting polynomial, the computation involves only the
orthogonal polynomials and the three associated vectors called the polynomial solution
vectors. The orthogonal polynomials are generated from the solution vectors whenever
it is necessary. Hence, the actual coefficients of the fitted polynomial are not required
for evaluation of derivatives of the computation of estimated values. However, if
desired, they may be printed by punching one (1) in the field. If this field is left blank
or contains a zero, the coefficients are not printed.
Evaluate Derivative Switch (cc 7-8)
This field is used to indicate whether derivatives are to be evaluated at a selected list
of points. If this field contains a zero or is blank, the derivatives are not computed.
If this field contains a one (1), each data point is examined to determine whether the
derivative computation indicator for that point is nonzero. When a nonzero indicator is
located for a particular value of x, the program evaluates the kth-order derivatives,
kth-order
where k is an integer punched in cc 9-10. The printout includes the 1, 2,
derivatives and the estimated value of y for that particular value of x.
Maximum Order Derivatives (cc 9-10)
If cc 7-8 contains a one (1), this field must be punched with an integer, k, which is less
than or equal to the order of the polynomial. It indicates the maximum order derivative
to be computed when a nonzero derivative computation indicator is located. The printout
includes all lower-order derivatives as well as the maximum.
34
Predicted Values (cc 11-12)
The value of y, as estimated, is printed for each data point, x; y is also estimated and
printed with the orthogonal polynomials unless the user does not elect to have them
printed (cc 31-32).
Punch Solution Vectors Switch (cc 13-14)
If this field contains a one (1), the polynomial solution vectors are punched in the
standard matrix format (section 2.1.4). If the field contains a zero or is left blank,
the solution vectors are not punched.
To maintain maximum numerical accuracy in computing the coefficients of the fitted
polynomial derivatives and estimated values, the orthogonal polynomials are used for
the computation. However, to conserve storage space, the orthogonal polynomials are
recomputed each time they are used. The polynomial solution vectors as functions of
the data points are used as parametric vectors in this computation to avoid making more
than one pass through the original data. In effect, the solution vectors are used
throughout the program to represent the coefficients of the fitted polynomials. Hence,
if the user expects to use the polynomial to compute additional values or derivatives,
the solution vectors should be punched out. In addition to the solution vectors, the first
output card includes the scaling constants required for evaluating y and derivatives at
x'. If scaling is not performed, this card is still punched, and read (but not used), under
input mode 3.
Variance Criterion (cc 17-26)
This field must be punched with a positive floating-point number of the form
.XXXXXXXXX. The number should include the decimal point, which may be placed
anywhere in the field. No blank columns are allowed.
The variance criterion is used to determine when the best fitting polynomial has been
computed. The process of fitting the points involves the computation of successively
higher-degree polynomials. As each degree computation is completed, a variance
criterion is developed. When the difference between any two successive variances is
less than the variance criterion punched in this field, the best fitting polynomial is
assumed to have the degree of the last polynomial computed. If, however, this condition
is not met before the maximum degree polynomial, as defined in cc 1-2, is satisfied,
the maximum degree is the degree used. If the user has no feeling for the magnitude of
this number, .01 may be used.
Transformation Switch (cc 27-28)
If this field is nonzero, a user-written transformation routine is called (see section
2.5.1). TRAN is called before any scaling that the user might elect to have the program
perform.
95
Scaling Switch (cc 29-30)
If this field is nonzero, x is scaled by a linear transformation to x' sach that the
elements of x are in the range (-2, 2). All calculations then deal with the data set
(x',y) (see section 2.4).
Polynomial and Residual Output (cc 31-32)
If this field is nonzero, only the coefficients of the polynomial y=a 0 + a ix +
are
re presented; the residuals, y-y*, y*, and . x ) are not listed.
Pi(
+ an xn
This option is helpful if solution vectors are punched, for later evaluation of y at new
points x*, when y at x is not desired.
Curve Fitting Type I Option Card Summary
Column
96
Meaning
1-2
Maximum degree of polynomial to be fitted
3-4
Input source
1 - Raw data input from card reader
2 - Raw data input from disk
5-6
Coefficients of fitted polynomial switch
0 - Do not print
1 - Print
7-8
Evaluate derivatives switch
0 - No derivatives
1 - Derivatives
9-10
Maximum order derivative
11-12
Predicted values
0 - Do not print
1 - Print
13-14
Polynomial solution vectors punch switch
0 - Do not punch
1 - Punch
15-16
Must be zero or blank
17-26
Variance criterion
27-28
Transformation switch (to user-written program)
0 - Do not call TRAN
1 - Call TRAN
29-30
Nonzero: Scale x into (-2, 2)
31-32
Nonzero: Do not print polynomials, predicted values, or residuals
TYPE II OPTION CARD
This type of option card is to be used whenever data is entered into the program with a
previously computed set of polynomial solution vectors.
Degree of Polynomial (cc 1-2)
This field is used to transmit the degree of the previously computed polynomial solution
vectors that are to be read by the program.
Input Type (cc 3-4)
This field must be punched with a three (3).
This program uses this number to read in the polynomial solution vectors that were
punched from a previous analysis. In addition to these solution vectors, necessary
scaling constants from the previous analysis are also read. It is necessary to keep all
punched output from analyses in the order in which it was punched, for later input.
In effect, the solution vectors represent the coefficients of the fitted polynomial. Hence,
this option is to be used when it is desired to use the fitted polynomial to compute
additional estimated values and/or to compute additional derivatives for points other
than those used in the initial analysis. The data points are read under the secondary
input type indicated in cc 15-16. If no data points are read for evaluation of y, only
coefficients are calculated. However, in this case, these have already been calculated
for the previous analysis. If the secondary input type is the card reader, previously
read input, which was placed on the disk, is destroyed.
Coefficients of Fitted Polynomial (cc 5-6)
In the determination of the best fitting polynomial, the computation involves only the
orthogonal polynomial and the three associated vectors called the polynomial solution
vectors. The orthogonal polynomials are generated from the solution vectors whenever
it is necessary. Hence, the actual coefficients of the fitted polynomial are not required
for evaluation of derivatives or the computation of estimated values. However, if
desired, they may be printed by punching a one (1) in this field. If this field is left
blank or contains a zero, the coefficients are not printed.
Evaluate Derivative Switch (cc 7-8)
This field is used to indicate whether derivatives are to be evaluated at a selected list
of points. If this field contains a zero or a blank, the derivatives are not computed.
If this field contains a one (1), each data point is examined to determine whether the
derivative computation indicator for that point is nonzero. When a nonzero indicator is
located for a particular value of x, the program evaluates the k th-order derivatives,
where k is an integer punched in cc 9-10. The printout includes the 1,2,
, kth-order
derivatives and the estimated values of y for that particular value of x.
97
Maximum Order Derivatives (cc 9-10)
If cc 7-8 contains a one (1), this field must be punched with an integer, k, which is less
than or equal to the order of the polynomial. It indicates the maximum order derivative to be
computed when a nonzero derivative computation indicator is located. The printout
includes all lower-order derivatives as well as the maximum. If the order is given as
zero, no derivatives are calculated.
Estimated Value Switch (cc 11-12)
This field is used to indicate whether estimated values are to be computed for the
values of x read in by the programs. If this field contains a zero or is left blank, the
estimated values are not computed. If a one (1) is punched, the estimated values are
computed.
Secondary Input Sources (cc 15-16)
This field must be punched with an integer one (1) or two (2).
If a one (1) is entered, the data points, followed by a negative identification card, are
read from the 1442 card reader, and onto the disk, destroying previously stored data.
If a two (2) is entered, the data points are read from the disk. If no data points are
entered, the format card is still required, and the first card succeeding the solution
vector deck is read as the new data card. If this card is blank, a y at x = zero is
evaluated.
If the user-written program TRAN was called, for initial analysis of the data, the user
should be sure that it is called to operate on the new data.
Curve Fitting Type II Option Card Summary
Column
98
Meaning
1-2
Degree of polynomial
3-4
Input type
3 - Polynomial solution vectors from card reader
5-6
Coefficients of fitted polynomial
0 - No
1 - Yes
7-8
Evaluate derivatives
0 - No
1 - Yes
9-10
Maximum order derivatives
Column
Meaning
11-12
Compute estimated values
0 - No
1 - Yes
13-14
Not used
15-16
Secondary input source
1 - Raw data from card reader
2 - Raw data from disk
17-26
Not used
27-28
Transformations
0 - Do not call TRAN
1 - Call TRAN
29-32
Not used (however, x is scaled if the x's were scaled for the
previous analysis)
2.4.3 Data Input
Raw data input to the program consists of a set of points (x, y) punched on cards, one
point to a card, with associated identification. The general form for this input can be
described in terms of individual fields for each item on the card.
Meaning
Field
Type
1
Integer (I)
2
Integer (1) Derivative computation indicator. The program computes
derivatives at any point specified in the data set. If the user
wishes to have a derivative of the polynomial evaluated at the
point punched on this card, this field should contain a one (1).
The order of derivative to be computed is specified on the
option card. If this field contains a zero (0), the derivative
is not evaluated at this point.
3
Floating
point (F)
The value of x. Any floating-point number is allowed. The
values of.x need not be equally spaced. Scaling of data may
be required.
Floating
point (F)
The value of y. Any floating-point number is allowed.
Card identification. Any numeric information that serves to
identify the point (x, . The number punched in this field
must be greater than zero.
99
The particular card columns for each field are arbitrary as long as all four fields are
present on the card.
Following the data deck, the user must include a card containing a negative integer in
the identification field. This card signals the program that no more data points are to
be processed.
2. 4. 4 Operating Instructions
A. Using the polynomial regression program when the total 1130 Statistical System
has not been stored on the disk
If the user wishes to load only the set of programs that allow this type of analysis,
the following programs must be compiled or assembled and stored on the disk. Each
deck begins with a card punched as
//FOR
and ends with an
*STORE
card.
The user should use a disk containing the 1130 Disk Monitor System, as described
in section 1.1. The following decks should be preceded by a cold start card, placed
in the card reader hopper, and the buttons IMMEDIATE STOP (console), RESET
(console), START (card reader), and PROGRAM LOAD (console) should be pressed.
A blank card should be placed after the last deck in the card reader hopper.
DECKS-LABELS: POLY-POLY; POL2-POL2; POLSQ-PLSQ;PCOEF-PCOF;
PFIT-P FIT; PDER-PDER; *F MAT-FMAT; *FMTRD-FMRD; *DATRD-DTRD;
*PRNTB-PRNB; *GMPYX-GMPY; *GDIVX-GDIV; TRAN-TRAN.
B. Execution from Disk
Once the component subroutines and main calling programs are on the disk, the
execution of a job requires the monitor control cards, program control cards, and
data cards to be placed in the card reader. The deck should be preceded by a cold
start card. To initiate processing, the buttons IMMEDIATE STOP and RESET
(console), START (card reader), and PROGRAM LOAD (console) should be pressed.
The order in which the cards are placed in the card reader for solution vector, disk,
or raw data input is shown in Figures 14, 15, and 16.
*Used in all four analysis types
100
/Optional
Blank Output
K
ota Deck
I
a
/Monitor
Control
(Option
(Job-Title
/Input/Output
Units
Variable
Format
/Negative
Identification
/
/
/
/
Figure 14. Orthogonal polynomial card order — card reader input
/Optional
Blank Output
/
/Monitor
(Option
Job-Title
Input/Output
Units
Control
Figure 15. Orthogonal polynomial card order — disk input
/
End of Data
../
(New Data
/Deck
/
Polynomial
Solution
Vectors */
/
/
Variable
Format
(Option
(Job-Title
Input/Output
Units
onitor
MControl
/
*Including scaling constants card
/
Figure 16. Orthogonal polynomial card order — solution vector input
101
2.4.5 Sample Problem
The data for this sample problem was taken, with permission, from page 213,
Statistical Theory in Research, by R. L. Anderson and T. A. Bancroft.
McGraw-Hill Book Company, Inc., New York, 1952.
INPUT
// XEQ POLY
02
*LOCALPOL2,POLSQ,PCOEF,PDER,PFIT
*LOCALPOLY,TRAN,DATRD,FMTRD,PRNTB
020200
1111
ORTH POLY (NO SCALING)
0201010102010100.010000000000000
(12,I1,1X,F2.0,F3.1)
011 01011
020 02071
031 03110
040 04126
051 05147
060 06199
071 07251
080 08239
091 09231
100 10236
111 11260
120 12246
-1
The numbers on this first card
are valid only when the user
elects to scale (see section
2.4.2). When scaling is not
performed, they reflect prior
core status and should be
ignored (that is, they can
take on any value).
PUNCHED OUTPUT
0.3636363E 00 0.1636363E 01 0
111124 1 1 0.6500000E 01 0.11"1666E 02 0.1772499E 02.
111124 1 2 0.6500000E 01 0.933332SE 01 0.2105245E 01
111124 1 3 0.6500000E 01 0.8678567E 01-0.2561686E 00
OUTPUT
102
// XFO POLY
0?
*LOCALPOL2oPOLSOOCOFF,PDER,PFIT
*LOCALPOLY.TRAN,DATRD.FMTRD,PRNTB
ORTH POLY (NO SCALING)
MAXIMUM DEGREE OF POLYNOMIAL
INPUT TYPE
POLYNOMIAL COEFFICIENTS
COMPUTE DERIVATIVES
ORDER OF DERIVATIVE
PREDICTED VALUES
PUNCH SOLUTION VECTORS
SECONDARY INPUT TY P E
VARIANCE CRITERION
TRANSFORMATION SWITCH
SCALING
IGNORE POLYNOMIAL OUTPUT
2
1
1
1
2
1
1
0
0.010000001
0
0
(I2.11,1X.F2.0.F3.1)
X = X' (NO TRANSFORMATION)
MAX DEGREE OF POLYNOMIAL REACHED. VARIANCE CRITERION NOT SATISFIED
103
ORTH POLY
(NO SCALING)
IDENTIFICATION
X,
-1
0.10000E
1
0.20000E
2
2
0.30000E
3
-3
4
4
0.40000E
5
0.50000E
-5
0.60000E
6
6
-7
0.70000E
7
0.80000E
8
P
-9
9
0.90000E
10
0.1D000E
IO
11
-11
0.11000E
0.12000E
12
17
JOB
Y
0.11000E
0.71000E
0.11000E
U.12600E
0.14700E
0.19900E
0.25100E
U.23930E
0.23100E
0.23600E
0.26000E
0.24600E
01
01
U1
01
01
01
01
01
01
02
02
U2
01
01
02
02
02
02
02
02
02
02
02
02
ORTHOGONAL POLYNOMIALS
Y•
Y-Y•
0.14497E 01 -0.34972E 00
0.61166E 01
0.98334E 00
0.10271E U2
0.72875E 00
0.13913E 02 -0.13134E 01
0.17043E 02 -0.23434E 01
0.19661E 02
0.23899E 00
0.21766E 02
0.33337E 01
0.23359E U2
0.54084E 00
0.24439E 02 -0.13397E 01
0.25007E 02 -0.14079E 01
0.25063E 02
0.93614E 00
0.24607E 02 -0.74081E-02
ALPHA
BETA
C
0
0.10000E
0.10000E
0.10000E
0.10000E
0.10000E
0.10000E
0.10000E
0.10000E
0.10000E
0.10000E
U.10000E
0.10000E
0.65000E
0.11916E
0.17724E
01
01
01
01
01
01
01
01
01
01
01
01
01
02
02
1
-0.55000E
-0.45000E
-0.35000E
-0.25000E
-0.15000E
-0.50000E
0.50000E
0.15UU0E
0.25000E
0.35000E
0.45000E
0.55000E
0.65000E
0.93333E
0.21052L
RAGE
1111
01
01
01
01
01
00
00
01
01
01
Ul
01
U1
01
01
2
0.18333E
0.83333E
0.33333E
-0.56866E
-0.96666E
-0.11666E
-0.11666E
-U.96666E
-0.56866E
0.33333E
0.83333E
0.16333E
0.65000E
0.86705E
-0.25616E
1
02
U1
OU
01
01
O.
02
ul
Ul
00
01
u2
01
V1
00
ANALYSIS OF VARIANCE
VARIATION SOURCE
SUM OF SOLARES
MEAN SOUARL
DEGREE 1 COMPONENT
RESIDUALS(DEGREE 1 REGR.(
1
10
0.63378E 03
0.11254E 03
0.63378E 03
0.11254E 02
DEGREE 2 CO M PONENT
RESIDUALS(DEGREE 2 REG.)
1
9
0•8 758 3E 02
0.24957E 02
0.87583E 02
0.27730E 01
00F•
ORTH POLY (NO SCALING)
JOB
1111
PAGE
2
COEFFICIENTS OF FITTED POLYNOMIAL
0
1
2
■0.3729547E 01
0.5435437E 01
■0.2561686E 00
JOB
ORTH POLY (NO SCALING)
IDENTIFICATION
-1
-3
-5
-7
104
Xi
1.00000
3.00000
5.00000
7.00000
Y.
1.44972
DERIV. ORDER
2
4.92309
-0.51233
2
3.89842
-0.51233
2
2.87375
- 0.51233
1
10.27124
17.04342
21.76624
-9
9.00000
24.43972
- 11
11.00000
25.06385
DERIV. VALUE
2
1.84907
- 0.51233
1
2
0.82440
- 0.51233
1
4:).20027
- 0.51233
1
2
1111
PAGE
3
JOB
ORTH POLY (NO SCALINGI
IDENTIFICATION
1
-1
2
3
4
5
6
7
8
9
10
11
12
2
-3
4
-5
6
-7
8
-9
10
12
PAGE
4
YP
X'
0.10000E
0.20000E
0.30000E
0.40000E
0.50000E
0.60000E
0.70000E
0.80000E
0.90000E
0.10000E
0.11000E
0.12000E
1111
01
01
01
01
01
01
01
01
01
02
02
02
0.11000E 01
0.71000E 01
o.110o0E 02
0.12600E 02
0.14700E 02
0.19900E 02
0.25100E 02
0.23900E 02
0.23100E n2
0.23600E 02
0.26000E 02
0.24600E 02
0.14497E
0.61166E
0.10271E
0.13913E
0.17043E
0.19661E
0.21766F
0.23359E
0.24439E
0.25007E
0.25063E
0.24607E
01 -0.34972E 00
01
0.98334E 00
02 0.72875E 00
02 0.13134E 01
02 ■0.23434E 01
02 0.23899E 00
02 0.33337E 01
02 0.54084E 00
02 -0.13397E 01
02 ■ 0.14079E 01
02 0.93614E 00
02 -0.74081E-02
JOS COMPLETED
INPUT
// XEQ POLY
02
*LOCALPOL2.POLSQ,PCOEFODERIPFIT
*LOCALPOLY,TRANIDATRO/FMTROORNT8
020200
1111
ORTH POLY (SCALING)
0201010102010100.010000000000100
(I2,11.1X,F2.01F3.1)
011 01011
020 02071
031 03110
040 04126
051 05147
060 06199
071 07251
080 08239
091 09231
100 10236
111 11260
120 12246
-1
PUNCHED OUTPUT
0.3636363E 00-0.2363636E 01 1
111124 1 1 0.3178914E-06 0.1575755E 01 0.1772499E 02
111124 1 2-0.1016575E-06 0.1234159E 01 0.5789424E 01
111124 1 3 0.3039385E-06 0.1147578E 01-0.1937265E 01
OUTPUT
// XEO POLY
02
•LOCALPOL2oPOLSOOCOEF•DEROFIT
*LOCALPOLY,TRANOATRO•FMTROORNTB
105
JOB
ORTH POLY (SCALING)
1111
PAGE
0
MAXIMUM DEGREE OF POLYNOMIAL
2
INPUT TYPE
1
POLYNOMIAL COEFFICIENTS
1
COM PUTE DERIVATIVES
1
ORDER OF DERIVATIVE
2
PREDICTED VALUES
1
PUNCH SOLUTION VECTORS
1
SECONDARY INPUT TYPE
0
0.010000001
VARIANCE CRITERION
0
TRANSFORMATION SWITCH
SCALING
tGNORE POLYNOMIAL OUTPUT
1
0
112.114,1X.F200.F3.11
THE X VALUES HAVE BEEN TRANSFORMED TO X' . ( 0.3634363E 001 • X t ( •0.2363636E 011.
MAX DEGREE OF POLYNOMIAL REACHED. VARIANCE CRITERION NOT SATISFIED
ORTH ono,(SCALING)
X,
I) ENTIFICATION
1
..0.19999E
2
2
..0.16363E
3
-0.12727E
4
4
-0.90909E
5
-5
-0.54545E
6
$0.18181E
7
0.18181E
8
0.54545E
9
0.90909E
10
10
0.12727E
11
-.11
0.16363E
12
12
0.20000E
01
01
01
00
00
00
00
00
00
01
01
01
JOB
0.11000E
0.71000E
0.11000E
0.12600E
0.14700E
0.19900E
0.25100E
0.23900E
0.23100E
0.23600E
0.26000E
0.24600E
01
01
02
02
02
02
U2
02
02
02
02
02
ORTHOGONAL POLYNOMIALS
Y.
'We.*
0.14497E 01
0.34974E 00
0.98334E 00
0.61166E 01
0.10271E 02
0.72875E 00
0.13913E 02 -0.13134E 01
0.17043E 02
0.23434E 01
0.19660E 02
0.23902E 00
0.21766E 02
0.33337E 01
0.23359E 02
0.54086E 00
0.24439E 02 -0.13397E 01
0.25007E 02
19.14079E 01
0.25063E 02
0.93614E 00
0.24607E 02 •.0.74310E..02
ALPHA
BETA
C
0
1.000000
1.000000
1.000000
1.000000
1.000000
1.000000
1.000000
1.000000
1.000000
1.000000
1.000000
1.000000
0.31789E .. 04
0.15757E 01
0.17724E 02
1111
PAGE
1
1
2
-2.000000
2.42424Z
1.636363
1.10190
1.272727
0.044076
0.909091
-0.545454
..1.278234
1.542697
13.181818
0.181817
1.542697
0.545454
..1.278235
-0.749310
0.909090
1.272726
0.044071
1.101928
1.636363
1.999999
2.424242
■0.10165E ■ 06
0.30393E - 06
0.12341E 01
0.11475E 01
0.578946 01 -0.19372E 01
REAnY THE PUNCH WITH PLANK CARDS AND PRESS START ON THE PUNCH AND CONSOLE. TURN CONSOLE SWITCH 15 ON.
ANALYSIS OF VARIANCE
VARIATION SOURCE
D.E.
DEGREE
1 COMPONENT
RESIDUALSIDEGREE
1 REGR.)
10
DEGREE
2 COMPONENT
RESIDUALSIDEGREE
2 REGR.I
9
1
1
SUM OF SQUARES
MEAN SQUARE
0.63378E 03
0.11254E 03
0.63378E 03
0.11254E 02
0.87584E 02
0.24956E 02
0.87584E 02
0 . 27729E 01
ORTH POLY (SCALING)
JOB
COEFFICIENTS OF FITTED POLYNOMIAL
0
1
2
106
0.2077764E 02
0.5789424E 01
- 0.1937265E 01
1111
PAGE
2
ORTH POLY (SCALING)
JOB
IDENTIFICATION
X'
Y.
DERIV. ORDER
-1.99999
1.44974
1
2
13.53848
-3.87453
-3
-1.27272
10.27124
1
2
10.72064
-3.87453
-5
0.54545
17.04340
1
7.90280
2
-3.87453
0.18181
-9
21.76622
0.90909
-11
24.43971
25.06386
1.63636
1
5.08496
2
..3.87453
1
2.26712
2
-3.87453
1
2
-0.55071
-3.87453
ORTH POLY (SCALING)
IDENTIFICATION
1
2
3
4
5
6
7
8
9
10
11
12
'1
2
-3
4
-5
6
-7
8
10
12
JOB
X'
0.19999E
'0.163638
-0.12727E
-0.909098
0.545458
'0.18181E
0.18181E
0.54545E
0.909098
0.12727E
0.16363E
0.20000E
Y.
01
01
01
00
00
00
00
00
00
01
01
01
0.11000E
0.71000E
0.11000E
0.12600E
0.147008
0.19900E
0.25100E
0.23900E
0.23100E
0.23600E
0.26000E
0.24600E
01
01
02
02
02
02
02
02
02
02
02
02
0.144978
0.61166E
0.10271E
0.13913E
0.17043E
0.19660E
0.21766E
0.23359E
0.24439E
0.25007E
0.25063E
0.24607E
01
01
02
02
02
02
02
02
02
0
02
02
PAGE
3
DERIV. VALUE
-1
-7
1111
1111
PAGE
4
Y-Y*
.• 0.34974E 00
0.98334E 00
0.72875E 00
-0.131348 01
-0.234348 01
0.23902E 00
0.33337E 01
0.540868 00
.. 0.13397E 01
(").140798 01
0.93614E 00
."0.74310802
JOP COMPLETED
INPUT
02
// XEQ POLY
*LOCALPOL2,POLSQ,PCOEF,PDER,PFIT
*LOCALPOLY,TRAN,DATRD,FMTRD,PRNTB
020200
ORTH POLY (NO SCALING, SOLUTION VECTOR INPUT)
1111
020301010201000100000000000000000
(12,I1,1X,F2.1,F3.0)
0.3636363E 00 0.1636363E 01 0
111124 1 1 0.6500000E 01 0.1191666E 02 0.1772499E 02
111124 1 2 0.6500000E 01 0.9333328E 01 0.2105245E 01
111124 1 3 0.6500000E 01 0.8678567E 01-0.2561686E 00
010 05000
020 55000
031 10000
These numbers were produced
-1
in the first sample problem, and
OUTPUT
as explained there, are not used.
However, the card is necessary.
107
// XEO POLY
02
*LOCALPOL2,POLSO0PCOEF0PDER•FIT
*LOCALPOLY,TRANOATRO,FMTRDORNTB
ORTH POLY (NO SCALING. SOLUTION VECTOR INPUT)
MAXIMUM DEGREE OF POLYNOMIAL
INPUT TYPE
POLYNOMIAL COEFFICIENTS
COMPUTE DERIVATIVES
JOB
1111
PAGE
0
JOB
1111
PAGE
1
2
3
1
1
ORDER OF DERIVATIVE
PREDICTED VALUES
PUNCH SOLUTION VECTORS
SECONDARY INPUT TYPE
2
1
0
1
VARIANCE CRITERION
TRANSFORMATION SWITCH
0.000000000
0
SCALING
IGNORE POLYNOMIAL OUTPUT
0
0
(12.11.1X.F2.111F3.0)
X
X) (NO TRANSFORMATION)
ORTH POLY (NO SCALING. SOLUTION VECTOR INPUT)
COEFFICIENTS OF FITTED POLYNOMIAL
0
-0.3729551E 01
1
0.5435436E 01
0.2561686E 00
2
ORTH POLY (NO SCALING. SOLUTION VECTOR INPUT)
IDENTIFICATION
X .Y*
-3
1.00000
1.44971
JOB
DERIV. ORDER
4.92309
-0.51233
1
ORTH POLY (NO SCALING. SOLUTION VECTOR INPUT)
1
2
3
JOB COMPLETED
108
1
2
-3
XI
0.50000E 00
0.55000E 01
0.10000E 01
Y
JOB
Y*
2
PAGE
DERIV. VALUE
2
IDENTIFICATION
1111
y- y*
0.00000E 00 -0.10758E 01 0.10758E 01
0.00000E 00 0.18416E 02 -0.18416E 02
0.00000E 00 0.14497E 01 -0•14497E 01
1111
PAGE
3
INPUT
// XEQ POLY
02
*LOCALPOL2,POLSO,PCOEF,PDER,PFIT
*LOCALPOLY,TRAN,OATRO,FMTRO,PRNTEI
020200
1111
ORTH POLY (SCALING, SOLUTION
020301010201000100000000000000000
(12,(1,1X,F2.1,F3.0)
0.3636363E 00-0.2363636E 01 1
111124 1 1 0.3178914E-06 0.1575755E
111124 1 2-0.1016575E-06 0.1234159E
111124 1 3 0.3039385E-06 0.1147578E
010 05000
020 55000
031 10000
VECTOR INPUT)
01 0.1772499E 02
01 0.5789424E 01
01-0.1937265E 01
OUTPUT
// XEO POLY
02
*LOCALPOL2sPOLSO,PCOEF,PDER+PFIT
*LOCALPOLY,TRANsDATROfFMTRD,PRNTB
ORTH POLY (SCALING, SOLUTION VECTOR INPUT)
MAXIMUM DEGREE OF POLYNOMIAL
INPUT TYPE
POLYNOMIAL COEFFICIENTS
COMPUTE DERIVATIVES
ORDER OF DERIVATIVE
PREDICTED VALUES
PUNCH SOLUTION VECTORS
SECONDARY INPUT TYPE
VARIANCE CRITERION
TRANSFORMATION SWITCH
JOB
1111
PAGE
0
2
3
1
1
2
1
0
1
0.000000000
0
SCALING
IGNORE POLYNOMIAL OUTPUT
0
0
II2.11.1X.F2.11F3.0)
THE X VALUES HAVF BEEN TRANSFORMED TO X . •( 0.3636363E 00/9X r (0.2363636E 01).
ORTH POLY (SCALING. SOLUTION VECTOR INPUT)
JOB
1111
PAGE
1
COEFFICIENTS OF FITTED POLYNOMIAL
0
1
2
0.2077764E 02
0.5789424E 01
-0.1937265E 01
109
JOB
ORTH POLY (SCALING. SOLUTION VECTOR INPUT)
X'
IDENTIFICATION
-1.99999
-3
Y•
1.44974
IDENTIFICATION
1
2
3
JOB COMPLETED
110
1
2
43
4 0.21818E 01
4 0.36363E 00
4 0.19999E 01
Y
1
13.53848
2
-3.87453
JOB
Y•
PAGE
2
DERIV. VALUE
DERIV. ORDER
ORTH POLY (SCALING. SOLUTION VECTOR INPUT)
X'
1111
Y4Y•
0.00000E 00 4 0.10758E 01
0.10758E 01
0.00000E 00 0.18416E 02 4 0.18416E 02
0.00000E 00 0.14497E 01 -0.14497E 01
1111
PAGE
3
2.5 GENERAL NOTES ON THE PROGRAMS
2.5.1 Transformations
This feature was added to aid users who are familiar with programming (see the manuals
1130 FORTRAN Language (C26-5933) and 1130 Disk Monitor System (C26-3750)) in adding
transformation capability to the system. Currently, a subroutine TRAN is included in the
package, and is called on option by each main program, subsequent to the reading of each
observation. The current routine returns to its calling program immediately.
In implementing such a subroutine, the following points should be considered:
1. In the regression and factor analysis programs, the observation (row) X is in
COMMON storage, and can be reached by use of the COMMON statement in the
user-written program. The row X contains one observation on X 1 ,
, XK , and
TRAN could be written using the row X as an argument.
2. In the orthogonal polynomial program, TRAN is called after each reading of x i , yi,
which are elements of vectors X, Y, in COMMON. TRAN could have arguments
x. and/or y., or could use the COMMON statement.
3. For the analysis of variance program, TRAN should include the argument DATA,
containing the observation.
4. If a large transformation program is prepared by the user, storage requirements
may call for the use of LOCAL monitor facilities.
5. Transformations that modify the number of variables in the observation require
modification to the program supplying and analyzing the data. Such modifications
require programming knowledge of the package. For example, if one originally
entered ten variables, and wished to transform the sum of four of them into one
column of the observation matrix, the sum should be placed in the column of one of
the original variables, and the program would have access to the resulting ten
variables for its analysis. In this specific instance, the program would exit because
of a singularity.
2.5.2 Notes on Correlation and E igen Analysis
The regression and factor analysis programs contain options that, in proper combination,
cause program termination when the correlation matrix and the latent roots and vectors
have been calculated. For example, the "no print" option (cc 25-26 of the option card)
used with the option for printing the correlation matrix gives this facility in the
regression program.
With the matrix input option to factor analysis, and using option 2 for the number of
factors, rotation option 0, and communality option 0, eigenvalues of matrices can be
obtained. However, the number of eigenvectors is limited to ten, the maximum number
of rotatable factors.
111
2.5.3 Punched Matrix Output
Matrix Number Dimension Name
1
NxN
Raw cross products matrix
2
NxN
Adjusted cross products matrix
3
NxN
Variance-covariance matrix
4
NxN
Correlation coefficients matrix
5
NxN
Characteristic vectors
6
NxK
Principal axis factor matrix
7
KxK
Orthogonal transformation matrix
8
NxK
Orthogonal factor matrix
9
KxK
Transformation to oblique reference structure matrix
10
NxK
Oblique reference vector structure matrix
11
KxK
Correlations among oblique reference vectors
12
NxK
Oblique reference vector pattern matrix
13
KxK
Correlations between reference vectors and primary
factors
14
NxK
Oblique primary factor structure matrix
15
KxK
Correlations among oblique primary factors
16
NxK
Oblique primary factor pattern matrix
17
NxK
Factor score regression coefficients
21
1x1
Number of cases
The following matrices include two or three vectors, which are punched as column
vectors where the column dimension indicates the number of elements in the vector.
Matrix Number Dimension
No. of Elements
on Each Card
22
Nx2
2
Raw sums, raw sums of squares
23
Nx2
2
Means, standard deviations
24
*N x 3
3
Alpha, Beta, C (orthogonal
polynomial solution vectors)
In the above: N = number of variables
K = number of rotated factors
*N = order of the polynomial
112
Meaning
2.5.4 Scaling
The programs in the 1130 Statistical System allow large data sets and a general input
format. Thus, there is a possibility that scaling will be necessary.
In regression and factor analysis, a pooling option is allowed that uses the raw cross
products matrix. If the number of observations is large, or if some observed variable
readings are quite large, some inaccuracies may become evident in this matrix.
Sometimes, scaling by use of the Format statement can aid in the solution of this
problem. In other situations, a transformation of the variable may help. It is also
possible that scaling should take place before data entry.
In orthogonal polynomials, if the order of the polynomial is high, and/or the range of
x is large, the elements of the polynomials will change rapidly in magnitude (as the
order of the polynomial increases) so as to even exceed the range of the floating-point
number, resulting in underflow or overflow. If data entered into this program is such
that this happens, as evidenced, for example, by the residuals, the user can elect to
transform or scale the dependent and/or independent variables. A useful option in this
program automatically scales the independent variable into a range such that the
magnitudes of the successive polynomial elements are approximately uniform.
In summary, the programs in this system are data-dependent, as is the case for many
computer programs. In some programs, definite accuracy characteristics can be
stated. In data-dependent routines, these statements are difficult to make.
113
CHAPTER 3: GENERAL FLOWCHARTS
The following charts (Figures 17, 18, 19) describe, generally, the programs in this
package. More detailed flowcharts, and listings, are available in the Systems Manual
for this package. The Systems Manual is not distributed with this program unless
specifically requested.
Stepwise Multiple Regression — System Flow
Analysis of Variance — System Flow
Read and Print
Control
Cards
i
Input
Source
Data
Input Source
Data
Form Cross
Products
I
Compute
Deviates and
Mean Squares
Need
Correl.
Matrix
I
Generate
ANOVA
Table
Form Subsidiary and Correlation
Matrices
I
/
Exit
N
Figure 17. Analysis of variance and stepwise multiple regression
114
NO
Figure 18. Curve fitting with orthogonal polynomials
115
Read and Pa/
Control
Cards
1
Figure 19. Factor analysis
116
CHAPTER 4: SAMPLE PROBLEM TIMING
The table below gives times for each sample problem, from the reading of the first
monitor card to the end of the output listing.
The
1132
Printer was used as the output device.
Problem
Time
(min:sec)
Regression analysis (card input)
6:02
Regression analysis (correlation matrix input)
3:04
Orthogonal polynomials (card input, no scaling)
2:35
Orthogonal polynomials (cards, scaling)
2:35
Orthogonal polynomials (solution vector input, no scaling)
2:00
Orthogonal polynomials (solution vector input, scaling)
2:00
Analysis of variance
2:14
Principal components analysis (card input)
5:48
Factor analysis (correlation matrix input)
3:45
117
CHAPTER 5: ERROR MESSAGES
Following is a list of error messages presented to the user on the (optional) printer or
the typewriter:
Common to All Programs
1. AN ILLEGAL CHARACTER HAS BEEN ENCOUNTERED IN COLUMN (N) OF THE
ABOVE FORMAT CARD. CHANGE CARD AND RERUN JOB.
Action: Correct the format card and rerun the job. See section 1.2 (Variable Format Card).
2. AN ILLEGAL CHARACTER HAS BEEN ENCOUNTERED IN APPROXIMATELY
COLUMN (N) OF THE ABOVE DATA CARD. CHANGE CARD AND RERUN JOB.
Action: The format card and/or the data card is in error. Correct the card(s) and
rerun the job. See sections 1.3(1) and 1.2 (Variable Format Card), and the
data input section pertaining to the particular analysis being run.
3. INVALID INPUT OPTION. JOB TERMINATED.
Action: Data input mode is not 01, 02, or 03. See the section discussing the option
card (input mode) for the particular analysis being run.
Common to Regression and Factor Analysis
4. CARD (ID) IS OUT OF SEQUENCE. RERUN JOB.
Action: Check sequence number of card, revise, and rerun job. See section 2.1.3
or 2.2.5.
Regression Messages
5. MEAN SQUARE NONPOSITIVE. JOB TERMINATED.
Action: A format specification error could have caused data to be converted
incorrectly, or an ill-conditioned matrix (for example, one with high
correlations between independent variables) could have caused inaccuracy
in the inversion of the correlation matrix or in the calculation of mean
squares.
6. NO MORE DEGREES OF FREEDOM. JOB TERMINATED.
Action: The number of parameters being estimated is larger than the number of
observations. Increase the number of observations, or accept a model with
fewer parameters.
7. NO MORE VARIABLES SATISFY THE VARIANCE CRITERION. JOB TERMINATED.
Action: Modify the variance criterion (section 2.1.2), or accept one of the models
produced.
118
9
tiUO333-1
O
0
International Business Machines Corporation
Data Processing Division
112 East Post Road, White Plains, N.Y. 10601
(USA Only)
IBM World Trade Corporation
821 United Nations Plaza, New York, New York 10017
(International)