>

>

Why Feature Scaling Works: Condition Number as the Key to Gradient Descent Convergence

Why Feature Scaling Works: Condition Number as the Key to Gradient Descent Convergence | RISE Research

Focus

Gradient Descent Optimization, Feature Scaling Methods, Hessian Conditioning

Motivation

Machine Learning Theory, Optimization Geometry, Neural Network Normalization

About the project

This paper provides a geometric explanation for why feature scaling accelerates gradient descent convergence, an effect that is universally recommended in machine learning practice but rarely given rigorous theoretical justification. The author uses the condition number kappa of the cost function's Hessian matrix as the unifying diagnostic quantity linking feature variance disparity to optimization difficulty, first establishing and empirically validating this link on synthetic functions with controlled condition numbers. The paper then compares four normalization methods (standardisation, min-max scaling, robust scaling, and PCA whitening) across varying dataset sizes and feature scale imbalances. Whitening is found to force kappa to its theoretical minimum of 1 regardless of the original scale ratio on clean data, but this advantage collapses under outlier contamination since it relies on an accurate covariance estimate; standardisation and min-max scaling offer more practical trade-offs but leave residual ill-conditioning when features are correlated, and min-max scaling is shown to be particularly sensitive to outliers at high variance ratios. To address this outlier sensitivity, the paper proposes and validates a two-stage pipeline combining robust scaling (which uses the outlier-resistant interquartile range) with whitening, achieving near-optimal conditioning even on noisy datasets. The paper also studies a complementary family of 'pure factor scaling' methods (KRCS, MRF, and a composite method) that leave the underlying data untouched and instead rescale the gradient itself using running second moments, showing these can outperform raw gradient descent without altering the data's condition number, demonstrating that ill-conditioning is only one of several sources of slow convergence. Finally, the paper connects these static preprocessing strategies to the dynamic normalization methods used inside neural networks (batch normalization and layer normalization), arguing both operate by continuously reducing the effective condition number during training rather than as a one-time preprocessing step.

Want to build a standout academic profile?

Interested in research mentorship?

Book a free call
Book a free call

Check out more projects

Emerging Clinical Strategies in Cholangiocarcinoma: A Review on Targeted Therapy and Immunotherapy

By :

Vanshika G.

View

From Rule-Based to Self-Improving Agents: The Evolution of AI Price Collusion

By :

Sanya S.

View

Within a single food category (snacks) on Blinkit, do products with nutrition marketing labels receive significantly different customer ratings than unlabelled products?

By :

Aahana B.

View

Emerging Clinical Strategies in Cholangiocarcinoma: A Review on Targeted Therapy and Immunotherapy

By :

Vanshika G.

View

From Rule-Based to Self-Improving Agents: The Evolution of AI Price Collusion

By :

Sanya S.

View

How to Apply

1.

Parent Consultation Call

2.

⁠Research Application Form

3.

⁠Profile Shortlisting

4.

⁠Program Onboarding

How to Apply

1.

Parent Consultation Call

2.

⁠Research Application Form

3.

⁠Profile Shortlisting

4.

⁠Program Onboarding

How to Apply

1.

Parent Consultation Call

2.

⁠Research Application Form

3.

⁠Profile Shortlisting

4.

⁠Program Onboarding

RISE Research Logo - Rise Global Education - Rise Research

+1 (650)-910-5964
admin@riseresearch.com

650 California St Fl 7, San Francisco, CA 94108

Copyright © 2025 RISE Research

All rights reserved.