What Chooses the First Principal Component Is Not the Data but the Units ── The same sample points at 55.6503 degrees in cm, 89.7345 in m, and 4.1934 in mm ── Not one value was changed; they were multiplied ── Standardising settles it at 45.0000 degrees, but standardising is itself a choice ── [Paper 375]
Abstract
Principal component analysis is described as extracting the main direction in the data. This paper asks whether the data alone fix that direction──the answer is that they do not: the units do. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──principal component analysis, the difference between covariance and correlation matrices, and standardisation are all standard. We do not build multivariate analysis──how many components to retain, and the relation to factor analysis, are not entered. They are named and no more. We use no real data──actual body measurements are not used. A sample from a bivariate normal with correlation 0.6 is generated. We build no estimation theory──sampling error in eigenvalues and tests on them are not the subject. We do not treat more than two variables──only the bivariate case is computed. The structure is the same with more. We offer no interpretation──the first component is not given a name such as “build”. Relation to earlier papers: Paper 358 showed a depth-of-field number fixed by a convention──there the convention was the circle of confusion; here it is the unit. Paper 356 showed quantities of the same dimension to be different quantities──here quantities of different dimensions are put in one space. Paper 300 showed whether two things share a root is decidable──by that test, “the structure of the data” and “the choice of scale” have distinct roots. Paper 372 showed an exponent governing the conclusion──here the scale governs it. What is added is giving the direction and explained fraction under four scalings, showing m and mm to be nearly orthogonal, showing standardisation settling at 45 degrees, showing the explained fraction running from 0.7899 to 1.0000, and putting the separator on covariance against correlation. First, a sample of heights and weights is generated and analysed with only the units changed (Section 2). Second, this is the core of the paper. It points at 55.6503 degrees in cm, 89.7345 in m, and 4.1934 in mm (Section 2). Third, in m and in mm the first component picks out two nearly orthogonal axes (Section 2). Fourth, the explained fraction also moves, reaching 1.0000 in m and 0.9898 in mm (Section 3). Fifth, standardising settles it at 45.0000 degrees with a fraction of 0.7899 (Section 4). Sixth, the separator is covariance matrix against correlation matrix, not the data (Section 5). Principal component analysis is described as extracting the main direction in the data──but that direction is not fixed by the data alone. PCA looks for the direction of greatest variance, and “greatest” is measured by the size of the numbers──yet the size of numbers changes with units: 170 cm is also 1.70 m and 1700 mm, one length producing variances that differ by 10^6. So on one sample the direction moves 85 degrees──holding weight in kg and changing only the height unit, the first component points at 55.6503 degrees in cm, 89.7345 in m, and 4.1934 in mm. The m and mm answers differ by 85.5411 degrees ── nearly orthogonal. 89.7345 is nearly the weight axis and 4.1934 nearly the height axis, so “the main variation is weight” and “it is height” swap on the choice of unit alone. The same sample, the same people, and the values were only multiplied. The explained fraction moves too──1.0000 in m and 0.9898 in mm. Both read as “the first component explains nearly everything”, and the axis being explained is different. A high fraction does not guarantee that the direction means anything: push the units far enough and it approaches 1 as closely as you like. The eigenvalues (98.318922 against 6312.309583) are not comparable either, since their units differ. Standardising settles it──with both variables standardised the direction is exactly 45.0000 degrees and the fraction 0.7899. The two variances become equal, and the eigenvectors of a correlation matrix are always at 45 and 135 degrees in two variables. The fraction matches (1+rho)/2 for rho=0.6, so the formula closes. One thing separates them──whether the covariance or the correlation matrix is used; not the data. Where units are meaningful, the covariance matrix; where they are mixed, the correlation matrix ── not “which is right” but “what should weigh more”. By the criterion of Paper 300 they have distinct roots: the structure comes from the sample, the scale from the analyst, and the result is their product. To be said honestly──standardising is itself a choice ── the decision to give every variable equal weight, which is not neutral. Where a high-variance variable ought to weigh more, the covariance matrix is right. The choice does not vanish; it changes place. One last thing──when one says “the data speak”, someone chose the scale in which they speak. Changing a unit alone turns the first component by 85 degrees, so that direction is not a property of the data alone. Report an analysis without its scale and it cannot be reproduced. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 主成分分析は「データの中の主要な方向」を取り出すと説明される。本稿が問うのはその方向がデータだけで決まるかである──答えは決まらない。単位が決めている。新しい定理も法則も主張しない。 本稿の射程(射程注記):新しい定理も法則も主張しない──主成分分析・共分散行列と相関行列の違い・標準化はすべて既知である。多変量解析を作らない──主成分の個数の決め方や、因子分析との関係には立ち入らない。名前を挙げるにとどめる。実データを扱わない──実際の身体測定は用いない。相関 0.6 の二変量正規から生成した標本を使う。推定論を作らない──固有値の標本誤差や検定は主題ではない。三変数以上を扱わない──二変数の場合だけを計算する。変数が増えても構造は同じである。解釈を与えない──第一主成分に「体格」などの名前を付けることはしない。既刊との関係:論文358 は被写界深度の数が約束で決まることを示した──そこでは約束が許容錯乱円だった。ここでは単位である。論文356 は次元が同じでも別の量であることを示した──ここでは次元の違う量を同じ空間に並べている。論文300 は同根か別根かが判定できることを示した──「データの構造」と「単位の選択」は、その基準で別根である。論文372 は指数が結論を支配することを示した──ここでは尺度が結論を支配する。加えたのは、四つの尺度で第一主成分の向きと寄与率を数に出したこと、m と mm でほぼ直交することを示したこと、標準化すると 45 度に落ち着くことを示したこと、寄与率が 0.7899 から 1.0000 まで動くことを示したこと、分離子を「共分散行列か相関行列か」に置いたことである。 第一に、身長と体重の標本を作り、単位だけを変えて主成分分析にかける(第2節)。 第二に、これが本稿の芯である。 cm で 55.6503 度、m で 89.7345 度、mm で 4.1934 度を向く(第2節)。 第三に、m と mm では、第一主成分がほぼ直交する二つの軸を指す(第2節)。 第四に、寄与率も動き、m で 1.0000、mm で 0.9898 になる(第3節)。 第五に、標準化すれば 45.0000 度・寄与率 0.7899 に落ち着く(第4節)。 第六に、分離子は「共分散行列か相関行列か」であって、データではない(第5節)。 主成分分析は「データの中の主要な方向」を取り出すと説明される──だがその方向は、データだけでは決まらない。主成分分析は分散の大きい方向を探し、「大きい」は数値の大きさで測られる──ところが数値の大きさは単位で変わる。170 cm は 1.70 m でもあり 1700 mm でもあって、同じ長さが 10^6 倍の分散の違いを生む。だから同じ標本で、向きが 85 度動く──体重を kg に固定したまま身長の単位だけを変えると、第一主成分は cm で 55.6503 度、m で 89.7345 度、mm で 4.1934 度を向く。 m と mm の差は 85.5411 度で、ほぼ直交する。89.7345 度はほぼ体重の軸、4.1934 度はほぼ身長の軸なので、「主要な変動は体重だ」と「身長だ」が単位の選択だけで入れ替わる。同じ標本、同じ人、同じ測定であり、数値は掛けただけである。寄与率も動く──m で 1.0000、mm で 0.9898。どちらも「第一主成分でほぼ全部説明できる」と読めるが、説明している軸が別である。高い寄与率は、その方向が意味をもつことを保証しない──単位を極端にすればいくらでも 1 に近づく。第一固有値そのもの(98.318922 と 6312.309583)も、単位が違えば単位が違うので比較できない。標準化すれば落ち着く──両方を標準化すると向きはちょうど 45.0000 度、寄与率は 0.7899 になる。二つの分散が等しくなるので、相関行列の固有ベクトルは二変数なら常に 45 度である。寄与率も (1+rho)/2 に対応しており、式が閉じている。分けているものは一つ──共分散行列を使うか相関行列を使うかであって、データではない。単位に意味がある場合は共分散行列、単位がばらばらなら相関行列で、「どちらが正しいか」ではなく「何を重く見たいか」である。論文300 の基準で別根であって、データの構造は標本から、尺度の選択は分析者から来る。結果はその積である。正直に言えば──標準化もまた一つの選択である。「全ての変数を同じ重みにする」という決定であって中立ではなく、分散の大きい変数を重く見たい場合にはむしろ共分散行列が正しい。選択は消えず、場所を変えるだけである。最後に一つ──「データが語る」と言うとき、語らせている尺度は誰かが選んでいる。単位を変えただけで第一主成分が 85 度回るのだから、その方向はデータだけの性質ではない。分析の結果を報告するときは、尺度も一緒に報告しなければ、再現できない。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。