Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This seems well done and well-researched. I appreciate the diagrams and art and references to the likes of Rubin and Pearl.

A few headings down:

> Causal inference provides us with tools that allow us to answer the question of why something happens.

This is not necessarily so.

Randomized controlled trials suffer from black box problems the same as models. This is clear enough when thinking about something like a tutoring program. Suppose I randomly assign a bunch of schools to learn algebra with curriculum X and the rest to continue business as usual.

Program X does better, so we infer the program has a causal impact on algebra learning.

However, we still do not know for sure why program X does better, only that it does better. This is important to inform how to take what works about the program and apply it to other circumstances, adapt it, and so on.

I suppose compared to a big data set, we have a better "why" answer to the variation between the outcome and the treatment. The difference being that we actually know the cause of the observed effect with a trial, whereas with correlational analyses we're not so sure. But that's a very deflationary view of "why." I don't mean to be too cynical here; we can always push "real" causality one more level down. For example, suppose we figure out the secret sauce to better algebra teaching relates to a specifical pedagogical practice. We can then say "but why does that practice work? what does it do in the brain?" So I don't want be too reductive.

But even gold standard RCTs don't always give us a "why?" answer. I remember attending a conference about a decade ago among causal inference-devoted social researchers specifically about "the black box" of causal inference as it pertains to RCTs.



This is why experimentation tries to focus on a singular change at a time -- in short, you're right, and we can focus on iteration to iteration to tease out fully causal factors.

But you know that, based on randomized assignment (or at least representative assignment school by school) an impact that A versus B determines.

So you do know E(Y | do(X), Z) to a degree, at least partially.


It’s like the old “5 Whys” technique for getting at the root of a problem. But as you suggest, there may not be a root: we can always push “real” causality one level down. Whys all the way down.

But that doesn’t mean we can’t answer the question of why something happens. A cause doesn’t cease to be cause just because it also has a cause.


It tells us "why something happens" in that we observe differences in how good people are at algebra, and the "why" is that some people took this class that causally improved their scores.


What specifically about the class improved the scores? It could be something intrinsic like a method that helps students remember things better. Or it could just be because the class was new the teachers were more involved versus the ones who are teaching the same old curriculum. These causes suggest different changes. One suggests everyone should teach X. The other suggests the solution needs to get teachers more involved.


Nobody said there couldn’t be follow-up questions after an RCT. But i disagree that the presence of additional whys disproves the goal of causal inference as proving a “why”


That sounds more like poor experimental design.


Right, what causal inference really gives you is a tool for specifying assumptions about a model of a data generating process and then estimating a parameter representing the effect size (and the uncertainty surrounding it), provided that your specified assumptions are approximately correct and your measurements are sufficiently accurate.

It never directly answers a "why" or "how" type question. You provide the why/how and then use data to estimate "by how much?"


The experiment didn't imply causality it implied correlation. As you point out much more is needed to establish causality.


Clarify your thinking about what it means to ask "why" here: https://plato.stanford.edu/entries/questions/#PraAppExpCon


Why is a philosophical question. We can only point to correlations and if time is relevant we can assume that future data can not change past data and only then conclude that causality has happened




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: