Nice.... Thanks guys...
Showing posts with label econometrics. Show all posts
Showing posts with label econometrics. Show all posts
Tuesday, June 17, 2025
Wednesday, October 26, 2016
Debugging
Neither programmer nor coder am I, but I am trying to teach my econometrics students how to fix problems when their R code fails. Curious, I googled debugging, and not surprisingly the word has a colorful, if contested history. Here's a picture of the very bug that started it all, in legend if not in fact.
Saturday, August 6, 2016
Skills beget skills
Why place a particular emphasis on early childhood education? There are two reasons. First, young kids' brains may be more plastic, in which case investments in education or good parenting would yield a higher return in terms of learning. Second, the ability to learn depends dynamically on the child's previously acquired capacity to learn. That is, skills beget skills. If so, early education earns a kind of "double dividend" by adding skills directly and facilitating later skill acquisition.
How important are these two effects? Not an easy question to answer, because we cannot directly observe or measure either investment in skills or the skills themselves. We do, however, have imperfect measures related to skills, such as test scores. These indicators may allow one to estimate the latent unobservables.
That's precisely the subject of this paper by Agostinelli and Wiswall, "Estimating the Technology of Children's Skill Formation." The dry title and dense methodology could be a little daunting, but the results are important. Here are my takeaways. First, identification of the latent variables and their effects is sensitive to modeling assumptions. Figuring out which assumptions are reasonable seems a high priority for future research. Second, under their preferred assumptions, they find the following, using data from the National Longitudinal Study of Youth (NLSY):
How important are these two effects? Not an easy question to answer, because we cannot directly observe or measure either investment in skills or the skills themselves. We do, however, have imperfect measures related to skills, such as test scores. These indicators may allow one to estimate the latent unobservables.
That's precisely the subject of this paper by Agostinelli and Wiswall, "Estimating the Technology of Children's Skill Formation." The dry title and dense methodology could be a little daunting, but the results are important. Here are my takeaways. First, identification of the latent variables and their effects is sensitive to modeling assumptions. Figuring out which assumptions are reasonable seems a high priority for future research. Second, under their preferred assumptions, they find the following, using data from the National Longitudinal Study of Youth (NLSY):
- A child's skills are strongly affected by both investments and pre-existing skills.
- Both effects are larger for younger kids. The results strongly favor early investment.
- There is some evidence that early investments have a bigger effect—and thus presumably bigger bang for the buck—for less-skilled kids. This result differs from some past findings which had suggested a reinforcing effect between skills and investment.
- Investment in skills is an increasing function of family income as well as the mother's cognitive and noncognitive skills—the latter having a particularly large impact. Noncognitive skills are measured using standard survey-based metrics conducted as part of the NLSY.
- Because investment is greater for children from advantaged backgrounds, "endogenous investment increases inequality in children’s skills."
Finally, the authors use their results to estimate the benefits and costs of an income transfer of $1000 to a child's family in terms of its impact on childhood skill development. The only benefit accounted for is the impact of skills on the child's future income. The net benefits are substantial, as shown in the table below.
It's tempting to read too much into this result, given the way it is presented. There is no attempt in the paper to show that the effect of income is causal. Rather, family income could be correlated with something else affecting investment in skills, such as neighborhood effects, or father's skills. So there is no evidence here that a simple money transfer would have these salutary effects. What they have demonstrated is that kids from disadvantaged backgrounds are at a very big disadvantage indeed in accumulating skills that affect life prospects in a big way. Given the dynamic of skill acquisition, figuring out how to level the playing field early in life is a compelling research and policy priority.
Labels:
econometrics,
economics,
education,
labor,
policy
Thursday, February 11, 2016
Teaching Undergrad Econometrics with R
Slides from my presentation at the February Bay Area R Users Group meetup. Obviously I long ago got over any kind of anxiety about being perceived as a nerd...
Labels:
data,
econometrics,
R,
statistics,
teaching
Saturday, September 12, 2015
Guide to R for SCU Economics Students, updated
Our latest version. Help yourself. You'll learn a lot.
Friday, August 14, 2015
Thursday, June 18, 2015
Guide to R for Econ Students
I have published a new online version of our Guide to R for SCU Economics Students. The Guide consists of hands-on tutorials, using examples from the first half of Stock and Watson's excellent Introduction to Econometrics (Pearson). I used R Markdown and published the chapters to RStudio's free RPubs platform.
Help yourself!
Help yourself!
Thursday, April 23, 2015
Yellow Pad Report
This week's Santa Clara Economics Department seminar was presented by Giovanni Peri, who took the Amtrak down from his home department at UC-Davis. He presented his latest paper on a topic he has been studying for some time now: the effect of immigration on native-born workers. The paper, with Mette Foged, analyzes an extraordinarily rich longitudinal data set of Danish workers.
As is the case in the United States, low-skilled workers are overrepresented among recent Danish immigrants. The most basic "Econ 101" analysis would predict that these low-skilled immigrants would compete with native-born low-skilled workers, increasing the supply and depressing the wage along the demand curve. Indeed, one of the most influential economists working on immigration effects, George Borjas, has a paper elaborating on precisely this claim, entitled "The Labor Demand Curve is Downward Sloping."
Well I'm sure Giovanni would agree that the demand curve slopes downward, but it turns out that immigration does not necessarily drive down the wages of low-skilled native workers. In fact, as his new paper shows, low-skilled Danes actually benefited from the waves of low-skilled refugees who settled in Denmark after 1995. The reason appears to be that as low-skilled foreigners filled jobs as manual laborers, many Danes who had held these positions were upgraded to new jobs that were complementary to the manual labor and actually paid a little better. For example, a construction laborer might have been upgraded to foreman to supervise the new foreign workers. Thus the effect of the shift in supply of low-skilled workers was more than offset by a shift in the demand for low-skilled native workers. This is a recurring theme of Giovanni's work: low-skilled workers are not homogeneous, and in particular native-born and foreign-born workers are not perfect substitutes.
The beauty of the paper is in the empirics. Studying immigration effects is notoriously challenging because of the endogeneity problem– determining the direction of causation. For example, suppose we observe immigrants flooding into a city or region, and wages rising at the same time. Can we conclude that immigrants caused the wages to rise? Or is it the reverse: that a growing regional economy, with increasing labor demand and rising wages, attracted the flow of immigrants to those employment opportunities?
Giovanni's solution to this tricky problem exploits some special features of the refugee flows to Denmark and some details of Danish refugee policy, along with careful application of modern panel econometrics. A fine paper cogently and enthusiastically presented, with interesting lessons for immigration policy.
As is the case in the United States, low-skilled workers are overrepresented among recent Danish immigrants. The most basic "Econ 101" analysis would predict that these low-skilled immigrants would compete with native-born low-skilled workers, increasing the supply and depressing the wage along the demand curve. Indeed, one of the most influential economists working on immigration effects, George Borjas, has a paper elaborating on precisely this claim, entitled "The Labor Demand Curve is Downward Sloping."
Well I'm sure Giovanni would agree that the demand curve slopes downward, but it turns out that immigration does not necessarily drive down the wages of low-skilled native workers. In fact, as his new paper shows, low-skilled Danes actually benefited from the waves of low-skilled refugees who settled in Denmark after 1995. The reason appears to be that as low-skilled foreigners filled jobs as manual laborers, many Danes who had held these positions were upgraded to new jobs that were complementary to the manual labor and actually paid a little better. For example, a construction laborer might have been upgraded to foreman to supervise the new foreign workers. Thus the effect of the shift in supply of low-skilled workers was more than offset by a shift in the demand for low-skilled native workers. This is a recurring theme of Giovanni's work: low-skilled workers are not homogeneous, and in particular native-born and foreign-born workers are not perfect substitutes.
The beauty of the paper is in the empirics. Studying immigration effects is notoriously challenging because of the endogeneity problem– determining the direction of causation. For example, suppose we observe immigrants flooding into a city or region, and wages rising at the same time. Can we conclude that immigrants caused the wages to rise? Or is it the reverse: that a growing regional economy, with increasing labor demand and rising wages, attracted the flow of immigrants to those employment opportunities?
Giovanni's solution to this tricky problem exploits some special features of the refugee flows to Denmark and some details of Danish refugee policy, along with careful application of modern panel econometrics. A fine paper cogently and enthusiastically presented, with interesting lessons for immigration policy.
Labels:
econometrics,
economics,
immigration,
labor
Tuesday, March 17, 2015
Today's little epiphany
The worst possible case of omitted variable bias in estimating the causal effect of Z on Y would be if Z had no direct causal effect on Y whatsoever, but was highly correlated with Y entirely because of Z's correlation with some other confounding variable X that causes Y. Then the regression coefficient on Z would be nothing but omitted variable bias! But what if you changed your mind and decided you were actually more interested in the causal effect of X on Y? Then Z, of course, would be a perfect instrumental variable!
Tuesday, September 9, 2014
Amazing R Markdown
I teach a new basic econometrics course using R, and I am quite proud of my Guide to R for SCU Economics Students (available here), which features a series of instructive tutorials with accompanying R scripts and data. I revise it and revise it and revise it, using one of humankind's most infuriating creations, Word. If I change the code, or a graphic, I have to run the R and print and paste the results into the script. Then, make sure Word has not gone and F-ed up the formatting, then print to pdf and upload.
But lo and behold: R Markdown. Simple text entry, intuitive formatting, embedded R code that will run and show the code and/or results in your document, formatted for optimal clarity. Saved automatically to HTML. Post and fuhgeddaboudit. Open source and free. Seems almost to have been designed with my needs specifically in mind. Outstanding. The sooner I can move my Guide to Markdown the better.
Bill Gates, you really are a good man. But I will gradually wean myself from your bloated annoying products, I swear I will.
But lo and behold: R Markdown. Simple text entry, intuitive formatting, embedded R code that will run and show the code and/or results in your document, formatted for optimal clarity. Saved automatically to HTML. Post and fuhgeddaboudit. Open source and free. Seems almost to have been designed with my needs specifically in mind. Outstanding. The sooner I can move my Guide to Markdown the better.
Bill Gates, you really are a good man. But I will gradually wean myself from your bloated annoying products, I swear I will.
Friday, September 5, 2014
Confirmation or falsification?
Many economists are trained to believe that when they do empirical work, they are– at least ideally– engaged in "Popperian" falsificationist methodology. That is, evidence can never prove something to be true, but it can prove something to be false. And using classical statistical methods, it's easy to convince yourself that this is what you are up to, because the whole enterprise involves seeing whether you can reject a hypothesis. But in this post, Andrew Gelman explains quite clearly why we are generally not falsificationists, but rather confirmationists... or at best that we "bounce" between the two. He goes on a bit, but here is the money section:
Deborah Mayo and I had a recent blog discussion that I think might be of general interest so I’m reproducing some of it here.
The general issue is how we think about research hypotheses and statistical evidence. Following Popper etc., I see two basic paradigms:
Confirmationist: You gather data and look for evidence in support of your research hypothesis. This could be done in various ways, but one standard approach is via statistical significance testing: the goal is to reject a null hypothesis, and then this rejection will supply evidence in favor of your preferred research hypothesis.
Falsificationist: You use your research hypothesis to make specific (probabilistic) predictions and then gather data and perform analyses with the goal of rejecting your hypothesis.
In confirmationist reasoning, a researcher starts with hypothesis A (for example, that the menstrual cycle is linked to sexual display), then as a way of confirming hypothesis A, the researcher comes up with null hypothesis B (for example, that there is a zero correlation between date during cycle and choice of clothing in some population). Data are found which reject B, and this is taken as evidence in support of A.
In falsificationist reasoning, it is the researcher’s actual hypothesis A that is put to the test.
How do these two forms of reasoning differ? In confirmationist reasoning, the research hypothesis of interest does not need to be stated with any precision. It is the null hypothesis that needs to be specified, because that is what is being rejected. In falsificationist reasoning, there is no null hypothesis, but the research hypothesis must be precise.
In our research we bounce
It is tempting to frame falsificationists as the Popperian good guys who are willing to test their own models and confirmationists as the bad guys (or, at best, as the naifs) who try to do research in an indirect way by shooting down straw-man null hypotheses.
And indeed I do see the confirmationist approach as having serious problems, most notably in the leap from “B is rejected” to “A is supported,” and also in various practical ways because the evidence against B isn’t always as clear as outside observers might think.
But it’s probably most accurate to say that each of us is sometimes a confirmationist and sometimes a falsificationist. In our research we bounce between confirmation and falsification.
Suppose you start with a vague research hypothesis (for example, that being exposed to TV political debates makes people more concerned about political polarization). This hypothesis can’t yet be falsified as it does not make precise predictions. But it seems natural to seek to confirm the hypothesis by gathering data to rule out various alternatives. At some point, though, if we really start to like this hypothesis, it makes sense to fill it out a bit, enough so that it can be tested.
In other settings it can make sense to check a model right away. In psychometrics, for example, or in various analyses of survey data, we start right away with regression-type models that make very specific predictions. If you start with a full probability model of your data and underlying phenomenon, it makes sense to try right away to falsify (and thus, improve) it.
Dominance of the falsificationist rhetoric
That said, Popper’s ideas are pretty dominant in how we think about scientific (and statistical) evidence. And it’s my impression that null hypothesis significance testing is generally understood as being part of a Popperian, falsificiationist approach to science.
So I think it’s worth emphasizing that, when a researcher is testing a null hypothesis that he or she does not believe, in order to supply evidence in favor of a preferred hypothesis, that this is confirmationist reasoning. It may well be good science (depending on the context) but it’s not falsificationist.P.S. for all you smug Bayesians: Read on. You're not off the hook.
Labels:
econometrics,
economics,
philosophy,
statistics
Tuesday, February 25, 2014
At Play in the Fields of Data
I continue reading / skimming my way through Gelman and Hill's Data Analysis Using Regression and Multilevel/Hierarchical Models. I very much like the tone of the book. It is practical... not doctrinaire. Plenty of examples. R code where you need it. Confidence intervals are plus or minus 2 standard errors, not 1.96 or whatever Student's t requires. Think about scales and units and log transformations... Don't just think about them: Try them out! Mess around with the data. Make some plots comparing your confidence intervals. Questions of causation are yet to come, but I anticipate that Gelman and Hill are not structural purists, nor identification cops. They want you to think about your problem, know your data, and especially be aware of other related results.
Everything has been comfortably familiar until the chapter on simulation of probability models. Toto, we're not in Kansas anymore... welcome to the land of Bayes.
Everything has been comfortably familiar until the chapter on simulation of probability models. Toto, we're not in Kansas anymore... welcome to the land of Bayes.
Thursday, February 20, 2014
How big is it?
A question that comes up (or that should come up!) in empirical research is: How big is the effect? How important? How does this effect compare with that one? Deirdre McCloskey calls it the question of "oomph." For example, in explaining variation in earnings across individuals, which has more oomph: differences in gender, in education, or in work experience?
To answer, we need estimates of the partial effects, but also a way to scale the units to make comparisons between apples and oranges. One conventional way to do this is to standardize regression coefficients by calculating the effect of a one-standard-deviation change in each variable. But as I realized while teaching this to my econometrics students this week, the comparison is tricky when some of your explanatory variables are qualitative (0-1), such as gender. What does it mean to predict the effect of a standard deviation of female-ness? (Yeah, OK, maybe something, but still...)
During my midterm today I started reading Gelman and Hill's Data Analysis Using Regression and Multilevel/Hierarchical Models (they really could have used a catchier title!). It's interesting to read a first-rate non-economist statistician on the techniques we economists use routinely. I have already gained one tip that helps with the problem at hand- the problem of oomph. Whereas standardized coefficients usually look at the effect of a one standard deviation change in X, Gelman recommends scaling regression coefficients by two s.d. Why? This makes comparison with the 0-1 effects of dummy variables more reasonable. When the mean of a dummy variable is around 0.5 (e.g., female), then one s.d. of it is also 0.5, so a 0-1 switch is 2 standard deviations of the dummy. And it's pretty close even when p = 0.2: s.d. = 0.4. Voila! I can compare the size of the effect of gender on earnings with the effect of experience or education.
Nifty! And practical! And easy... And did I mention that all their examples are cleanly coded in R? That's especially handy for ECON 41/42 at Santa Clara University. It's a fat book, and it gets harder, but I'll keep reading.
To answer, we need estimates of the partial effects, but also a way to scale the units to make comparisons between apples and oranges. One conventional way to do this is to standardize regression coefficients by calculating the effect of a one-standard-deviation change in each variable. But as I realized while teaching this to my econometrics students this week, the comparison is tricky when some of your explanatory variables are qualitative (0-1), such as gender. What does it mean to predict the effect of a standard deviation of female-ness? (Yeah, OK, maybe something, but still...)
During my midterm today I started reading Gelman and Hill's Data Analysis Using Regression and Multilevel/Hierarchical Models (they really could have used a catchier title!). It's interesting to read a first-rate non-economist statistician on the techniques we economists use routinely. I have already gained one tip that helps with the problem at hand- the problem of oomph. Whereas standardized coefficients usually look at the effect of a one standard deviation change in X, Gelman recommends scaling regression coefficients by two s.d. Why? This makes comparison with the 0-1 effects of dummy variables more reasonable. When the mean of a dummy variable is around 0.5 (e.g., female), then one s.d. of it is also 0.5, so a 0-1 switch is 2 standard deviations of the dummy. And it's pretty close even when p = 0.2: s.d. = 0.4. Voila! I can compare the size of the effect of gender on earnings with the effect of experience or education.
Nifty! And practical! And easy... And did I mention that all their examples are cleanly coded in R? That's especially handy for ECON 41/42 at Santa Clara University. It's a fat book, and it gets harder, but I'll keep reading.
Tuesday, April 16, 2013
How not to Excel at economic research
Turns out that that the widely cited work of Reinhart and Rogoff alleging an adverse impact of high national debt on economic growth suffered from some serious errors, including an incorrect cell range reference in an Excel spreadsheet.
I won't comment here on the substance of the debate, nor on whether R&R have provided an adequate response to their critics. But I would like to draw attention to one lesson I draw from the whole kerfuffle, and that is that Excel spreadsheets, whatever their virtues, are ill-suited to careful data work. R&R's mistake is a case in point. It is simply too easy to introduce errors in cell ranges and references in Excel, and way too hard to replicate results. This is not to say that scripts in SAS, Stata, or R are foolproof... far from it. But this kind of mistake, where part of the sample is left out inadvertently, would be quite unlikely, I think. Furthermore, Excel is ill-suited to sensitivity and specification checks. In smallish samples like the ones R&R are dealing with, sensitivity checks are essential.
I won't comment here on the substance of the debate, nor on whether R&R have provided an adequate response to their critics. But I would like to draw attention to one lesson I draw from the whole kerfuffle, and that is that Excel spreadsheets, whatever their virtues, are ill-suited to careful data work. R&R's mistake is a case in point. It is simply too easy to introduce errors in cell ranges and references in Excel, and way too hard to replicate results. This is not to say that scripts in SAS, Stata, or R are foolproof... far from it. But this kind of mistake, where part of the sample is left out inadvertently, would be quite unlikely, I think. Furthermore, Excel is ill-suited to sensitivity and specification checks. In smallish samples like the ones R&R are dealing with, sensitivity checks are essential.
Friday, March 8, 2013
My new course,
Data Analysis and Econometrics, is bound to become the hottest class on campus. Nothing impresses that special someone like murmuring in her/his ear, "Baby, I know how to program R to correct regression standard errors for heteroskedasticity."
Tuesday, September 25, 2012
Saturday, July 28, 2012
Stats R Us, Part 1
Here's my initial blog post on teaching intro econometrics with R, posted on our Technology in Teaching blog.
Labels:
econometrics,
statistics,
teaching,
technology
Subscribe to:
Posts (Atom)
