Skip to content

What Chooses the First Principal Component Is Not the Data but the Units ── The same sample points at 55.6503 degrees in cm, 89.7345 in m, and 4.1934 in mm ── Not one value was changed; they were multiplied ── Standardising settles it at 45.0000 degrees, but standardising is itself a choice ── [Paper 375]

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Principal component analysis is described as extracting the main direction in the data. This paper asks whether the data alone fix that direction──the answer is that they do not: the units do. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──principal component analysis, the difference between covariance and correlation matrices, and standardisation are all standard. We do not build multivariate analysis──how many components to retain, and the relation to factor analysis, are not entered. They are named and no more. We use no real data──actual body measurements are not used. A sample from a bivariate normal with correlation 0.6 is generated. We build no estimation theory──sampling error in eigenvalues and tests on them are not the subject. We do not treat more than two variables──only the bivariate case is computed. The structure is the same with more. We offer no interpretation──the first component is not given a name such as “build”. Relation to earlier papers: Paper 358 showed a depth-of-field number fixed by a convention──there the convention was the circle of confusion; here it is the unit. Paper 356 showed quantities of the same dimension to be different quantities──here quantities of different dimensions are put in one space. Paper 300 showed whether two things share a root is decidable──by that test, “the structure of the data” and “the choice of scale” have distinct roots. Paper 372 showed an exponent governing the conclusion──here the scale governs it. What is added is giving the direction and explained fraction under four scalings, showing m and mm to be nearly orthogonal, showing standardisation settling at 45 degrees, showing the explained fraction running from 0.7899 to 1.0000, and putting the separator on covariance against correlation. First, a sample of heights and weights is generated and analysed with only the units changed (Section 2). Second, this is the core of the paper. It points at 55.6503 degrees in cm, 89.7345 in m, and 4.1934 in mm (Section 2). Third, in m and in mm the first component picks out two nearly orthogonal axes (Section 2). Fourth, the explained fraction also moves, reaching 1.0000 in m and 0.9898 in mm (Section 3). Fifth, standardising settles it at 45.0000 degrees with a fraction of 0.7899 (Section 4). Sixth, the separator is covariance matrix against correlation matrix, not the data (Section 5). Principal component analysis is described as extracting the main direction in the data──but that direction is not fixed by the data alone. PCA looks for the direction of greatest variance, and “greatest” is measured by the size of the numbers──yet the size of numbers changes with units: 170 cm is also 1.70 m and 1700 mm, one length producing variances that differ by 10^6. So on one sample the direction moves 85 degrees──holding weight in kg and changing only the height unit, the first component points at 55.6503 degrees in cm, 89.7345 in m, and 4.1934 in mm. The m and mm answers differ by 85.5411 degrees ── nearly orthogonal. 89.7345 is nearly the weight axis and 4.1934 nearly the height axis, so “the main variation is weight” and “it is height” swap on the choice of unit alone. The same sample, the same people, and the values were only multiplied. The explained fraction moves too──1.0000 in m and 0.9898 in mm. Both read as “the first component explains nearly everything”, and the axis being explained is different. A high fraction does not guarantee that the direction means anything: push the units far enough and it approaches 1 as closely as you like. The eigenvalues (98.318922 against 6312.309583) are not comparable either, since their units differ. Standardising settles it──with both variables standardised the direction is exactly 45.0000 degrees and the fraction 0.7899. The two variances become equal, and the eigenvectors of a correlation matrix are always at 45 and 135 degrees in two variables. The fraction matches (1+rho)/2 for rho=0.6, so the formula closes. One thing separates them──whether the covariance or the correlation matrix is used; not the data. Where units are meaningful, the covariance matrix; where they are mixed, the correlation matrix ── not “which is right” but “what should weigh more”. By the criterion of Paper 300 they have distinct roots: the structure comes from the sample, the scale from the analyst, and the result is their product. To be said honestly──standardising is itself a choice ── the decision to give every variable equal weight, which is not neutral. Where a high-variance variable ought to weigh more, the covariance matrix is right. The choice does not vanish; it changes place. One last thing──when one says “the data speak”, someone chose the scale in which they speak. Changing a unit alone turns the first component by 85 degrees, so that direction is not a property of the data alone. Report an analysis without its scale and it cannot be reproduced. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 主成分分析は「データの中の主要な方向」を取り出すと説明される。本稿が問うのはその方向がデータだけで決まるかである──答えは決まらない。単位が決めている。新しい定理も法則も主張しない。 本稿の射程(射程注記):新しい定理も法則も主張しない──主成分分析・共分散行列と相関行列の違い・標準化はすべて既知である。多変量解析を作らない──主成分の個数の決め方や、因子分析との関係には立ち入らない。名前を挙げるにとどめる。実データを扱わない──実際の身体測定は用いない。相関 0.6 の二変量正規から生成した標本を使う。推定論を作らない──固有値の標本誤差や検定は主題ではない。三変数以上を扱わない──二変数の場合だけを計算する。変数が増えても構造は同じである。解釈を与えない──第一主成分に「体格」などの名前を付けることはしない。既刊との関係:論文358 は被写界深度の数が約束で決まることを示した──そこでは約束が許容錯乱円だった。ここでは単位である。論文356 は次元が同じでも別の量であることを示した──ここでは次元の違う量を同じ空間に並べている。論文300 は同根か別根かが判定できることを示した──「データの構造」と「単位の選択」は、その基準で別根である。論文372 は指数が結論を支配することを示した──ここでは尺度が結論を支配する。加えたのは、四つの尺度で第一主成分の向きと寄与率を数に出したこと、m と mm でほぼ直交することを示したこと、標準化すると 45 度に落ち着くことを示したこと、寄与率が 0.7899 から 1.0000 まで動くことを示したこと、分離子を「共分散行列か相関行列か」に置いたことである。 第一に、身長と体重の標本を作り、単位だけを変えて主成分分析にかける(第2節)。 第二に、これが本稿の芯である。 cm で 55.6503 度、m で 89.7345 度、mm で 4.1934 度を向く(第2節)。 第三に、m と mm では、第一主成分がほぼ直交する二つの軸を指す(第2節)。 第四に、寄与率も動き、m で 1.0000、mm で 0.9898 になる(第3節)。 第五に、標準化すれば 45.0000 度・寄与率 0.7899 に落ち着く(第4節)。 第六に、分離子は「共分散行列か相関行列か」であって、データではない(第5節)。 主成分分析は「データの中の主要な方向」を取り出すと説明される──だがその方向は、データだけでは決まらない。主成分分析は分散の大きい方向を探し、「大きい」は数値の大きさで測られる──ところが数値の大きさは単位で変わる。170 cm は 1.70 m でもあり 1700 mm でもあって、同じ長さが 10^6 倍の分散の違いを生む。だから同じ標本で、向きが 85 度動く──体重を kg に固定したまま身長の単位だけを変えると、第一主成分は cm で 55.6503 度、m で 89.7345 度、mm で 4.1934 度を向く。 m と mm の差は 85.5411 度で、ほぼ直交する。89.7345 度はほぼ体重の軸、4.1934 度はほぼ身長の軸なので、「主要な変動は体重だ」と「身長だ」が単位の選択だけで入れ替わる。同じ標本、同じ人、同じ測定であり、数値は掛けただけである。寄与率も動く──m で 1.0000、mm で 0.9898。どちらも「第一主成分でほぼ全部説明できる」と読めるが、説明している軸が別である。高い寄与率は、その方向が意味をもつことを保証しない──単位を極端にすればいくらでも 1 に近づく。第一固有値そのもの(98.318922 と 6312.309583)も、単位が違えば単位が違うので比較できない。標準化すれば落ち着く──両方を標準化すると向きはちょうど 45.0000 度、寄与率は 0.7899 になる。二つの分散が等しくなるので、相関行列の固有ベクトルは二変数なら常に 45 度である。寄与率も (1+rho)/2 に対応しており、式が閉じている。分けているものは一つ──共分散行列を使うか相関行列を使うかであって、データではない。単位に意味がある場合は共分散行列、単位がばらばらなら相関行列で、「どちらが正しいか」ではなく「何を重く見たいか」である。論文300 の基準で別根であって、データの構造は標本から、尺度の選択は分析者から来る。結果はその積である。正直に言えば──標準化もまた一つの選択である。「全ての変数を同じ重みにする」という決定であって中立ではなく、分散の大きい変数を重く見たい場合にはむしろ共分散行列が正しい。選択は消えず、場所を変えるだけである。最後に一つ──「データが語る」と言うとき、語らせている尺度は誰かが選んでいる。単位を変えただけで第一主成分が 85 度回るのだから、その方向はデータだけの性質ではない。分析の結果を報告するときは、尺度も一緒に報告しなければ、再現できない。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

View source

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Conference Open access Dec 2013

Affordable and Energy-Efficient Cloud Computing Clusters: The Bolzano Raspberry Pi Cloud Cluster Experiment

The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.

P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al. · 110 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6

Related blog posts

MIT News · Artificial Intelligence Sep 14, 2026

New method enables AI for safety-critical situations

The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.

GPT-Lab Sep 10, 2026

Responsible AI Must Consider Its Afterlife

AI may appear weightless, but every model depends on physical infrastructure. To understand responsible AI, we need to look beyond algorithms and consider the entire lifecycle of the hardware behind them. The post Responsible AI Must Consider Its Afterlife appeared first on GPT-Lab.

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.