← ireadpaper · 顶刊中的公共政策研究

预测收入分布的简约方法

A parsimonious approach to predicting income distributions
Journal of Development Economics · 2025 · [{"name": "Daniel Gerszon Mahler", "affiliation": []}, {"name": "Marta Schoch", "affiliation": []}, {"name": "Christoph Lakner", "affiliation": []}, {"name": "Minh Cong Nguyen", "affiliation": []}]

中文摘要

本文开发了一种方法,利用一个包含少量国家层面变量的简单回归,预测世界所有国家可比的收入与消费分布。为拟合该模型,分析使用了来自世界银行贫困与不平等平台、涵盖168个国家的约2,000个住户调查分布数据。研究使用了来自多个数据库的1,000多个经济、人口和遥感预测变量来检验模型。最终选定的模型在样本外准确性、简约性以及可适用国家占比之间取得了平衡。本文发现,一个仅依赖人均国内生产总值、5岁以下儿童死亡率、预期寿命和农村人口占比的简约模型,其准确性与联合使用1,000个指标的复杂机器学习模型几乎相同。这一组与人类发展相关的基础指标解释了收入分布的大部分跨国差异,即使在数据极度匮乏的国家,也能促进分布分析。

Abstract

This paper develops a method to predict comparable income and consumption distributions for all countries in the world from a simple regression with a handful of country-level variables. To fit the model, the analysis uses around 2,000 distributions from household surveys covering 168 countries from the World Bank's Poverty and Inequality Platform. More than 1,000 economic, demographic, and remote sensing predictors from multiple databases are used to test the models. A model is selected that balances out-of-sample accuracy, simplicity, and the share of countries for which it can be applied. The paper finds that a parsimonious model relying on gross domestic product per capita, under-5 mortality rate, life expectancy, and rural population share gives almost the same accuracy as a complex machine learning model using 1,000 indicators jointly. This small set of basic indicators related to human development explains most cross-country variation in income distributions and can facilitate distributional analysis even in countries with extreme data deprivation.
在 ireadpaper 查看全部 →