Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Monday, March 9, 2026

The world's deadliest animals...

... according to Our World in Data. Mosquitoes come in first, and humans second. Snakes are a distant third, followed by dogs. All of that is plausible to me. These guys do good work, and I trust their research, but I was surprised that neither fleas nor the rodents that carry them made the list. Surely the plague, tamed though it may be, still carries off more humans than the sharks that account for only six fatalities per year. And ticks? Nada? 

(Image: Wikipedia)


Thursday, October 13, 2022

Congrats, Steve Ruggles!

Recipient of a 2022 MacArthur "genius" grant– unexpected but richly deserved. Where would I be without IPUMS

Thursday, February 11, 2016

Teaching Undergrad Econometrics with R

Slides from my presentation at the February Bay Area R Users Group meetup. Obviously I long ago got over any kind of anxiety about being perceived as a nerd...

Saturday, December 5, 2015

Trump's language

This NYT article on the Donald's verbal demagoguery is interesting, based on some text analysis of all of his public utterances over the past week... but where are the tables and graphs? And more importantly, where are the comparisons with the rest of the Republican field? Is his tendency toward "us vs. them" really that different from Ted Cruz or the rest? Data analysis is all about comparisons.

Friday, November 6, 2015

"Household surveys in crisis"

This article summarizes some highly important work by Bruce Meyer and colleagues on the ongoing deterioration of the reliability of household survey data, and its implications for measuring important social and economic outcomes. Measurement is a dry and boring topic to many, but anyone who has relied on Census-type surveys to analyze social trends has to be quite concerned.

The Current Population Survey (CPS) is among the most important and widely used of these surveys. Conducted by the Census Bureau and the Bureau of Labor Statistics, it is the source of our monthly information on unemployment rates, as well as a standard source for income and poverty measures on an annual basis.

Our understanding of poverty as well as the effectiveness of anti-poverty policies rests critically on having good measures of the various sources of cash and non-cash income and assistance received by low-income Americans. As Meyer and Nikolas Mittag show in a new working paper, respondents to the CPS often fail to report receipt of various government transfers, and when they do they frequently under-report the amount. Hence the bias is strongly in the direction of over-estimating poverty, and under-estimating the poverty-reducing effect of transfer programs.

How can we tell what people are actually receiving? Meyer and Mittag link the CPS data for New York State to administrative records of the government agencies that are actually writing the checks. They find, for example, that for people reporting incomes at less than half the official poverty line, the missing transfers are worth a little more than the entire amount of their reported cash income. As a consequence of such measurement errors, "the poverty reducing effect of all programs together is nearly doubled" when transfers are fully accounted for.

The most poorly measured form of assistance in their New York data is housing subsidies. One reason for this may be that when these payments are made directly to the landlord, survey respondents may not be aware of the actual amount, and underestimate it. The CPS tries to impute rental assistance when it is not provided by respondents, but the imputation procedures are highly inaccurate, and downward-biased. Survey responses are incomplete and biased for other transfers as well, perhaps due to survey "fatigue", or stigma, or something else.

So there's bad news, but also some good news. The bad news is that these workhorse surveys are not nearly as reliable as we would have liked to believe, not through negligence or malfeasance on the part of the officials who run them, but because it is very difficult to elicit accurate responses. The good news is that administrative and other data sources may provide a more accurate picture. And better news still is the finding that thanks to government transfer programs, the social safety net is actually a lot more effective than we had thought.

Sunday, May 10, 2015

Why we love R...

Kieran Healy scrapes the UK election returns data from a BBC web site, cleans up the data, and maps it. Insightful, beautiful. All in R. I confess, Stata is my go-to statistical package. But could you do what Kieran did using Stata? Doubt it.


Monday, August 18, 2014

As a frequent data janitor myself...

... I feel their pain.
Data scientists, according to interviews and expert estimates, spend from 50 percent to 80 percent of their time mired in this more mundane labor of collecting and preparing unruly digital data, before it can be explored for useful nuggets.
The worst is when you get it all cleaned up and there are no useful nuggets to be found...

Thursday, May 17, 2012

Ignorance is bliss

"Whites Account for Under Half of Births in U.S."
"Annual Census at Risk in House Budget Bill"
Ah, the workings of the Republican mind... If we don't know they are there, maybe they aren't really there...!