I might have missed it in this rather long article, but I think the whole debate could use some people more that have read Karl Popper and are familiar with Positivism vs Falsifiability. I would not be surprised if all this amounts to is that you can of course "replicate" results if you try often enough, and ignore the falsifying results when it did not work.
I mean, of course the big problem is that applying strict falsifiability to social sciences does not work. You can always find a counter-example, the problems studied do not work like that. It is hard to reconcile this, but I think that is the heart of the replication crisis (together with some bad statistics). But (having done a PhD in a somewhat related field) I see only two options:
1. We need a complete overhaul of how experiments and study results are published, so that observers can see the failed results and we can try to assess how often a theory hold up
2. We have to limit non-hard sciences to questions that are not ambiguous, where one well-done negative result really shows a theory is wrong.
Of course, there is a third option: Just continue as it is now and ignore all results that seem unlikely, because given how those fields work, they are most likely wrong.
The problem with 2, is that psychology is happy to acknowledge that many of its results are context dependent [0]. The entire field of cross-cultural psychology is founded on the understanding that psychological findings may not generalize across cultures. While this suggest that there are likely no "natural laws" in psychology, it does nothing to dismiss the importance and usefulness of the non-natural laws we do find. Given this framework, falsifiability still applies, but may only suggest refinement of the conditions of our study.
Your third option seems to require the pre-empirical determination of which results of "unlikely" and instead choosing to believe only in the studies that confirm what we already intuitively believe.
Having said that, I completely agree with your first point. Replication studies alongside the increased visibility of studies that fail to reject the null hypothesis are absolutely essential.
> The problem with 2, is that psychology is happy to acknowledge that many of its results are context dependent
Yes. But that's only a specific kind of psychology. In my master I had a psychology professor (who came from the physics department originally) who clearly said all of that is bogus, and the only thing psychology should do are clear-cut studies like finding just noticeable differences. In his case, he was studying how the brain processes visual input that way. Very interesting, and worked well for him.
You simply need to declare studies before you preform them and you avoid the kind of bias you are talking about. The core issue is not the method, but poor incentives leading to people gaming the system instead of doing science. No system is going to survive most people trying to break it.
On the whole the 'soft' sciences are not actually science right now. They are philosophy playing dress up. Sure, you get a few honest people trying to do real research, but mostly not.
PS: I remember seeing the same problem on the small scale in the physical sciences. One professor tried to weigh air by weighing a bag with and without air. Which did not work for obvious reasons. But, blow on it adding enough saliva and eventually you can finagle the number you want. I was kind of hoping he was going to eventually say, see don't do this and then give details, but nope he felt like he got the results he wanted.
It's a common and disingenuous criticism that registered reports disallow discovery. They don't prevent you from doing exploratory analysis, they simply require you to display it as exploratory rather than pretending that was your hypothesis all along.
Yeah, I have very little respect for this criticism.
In fact, the end of the Slate article gives a perfect example of how things should work. Bem did a large, preregistered psi study, found no effect, but noticed an interesting positive result he hadn't registered for. So he published the negative result, mentioned the positive thread as a site for further research, and is now putting together a preregistered study on that basis.
This seems totally above board. (And again, Bem outperforms the standards of real fields.) There's nothing wrong with noticing something suggestive in a registered dataset, or even gathering exploratory data directly. You just have to confirm what you find with preregistration.
The problem with small changes is they are going to benefit or hurt each party. If you say A || B || C || D then party X will like A & D, and party Y will like B & C but there is no way to move forward.
Just look at giving DC a house seat when it's larger than other states who get not only a seat in the house but 2 in the senate. Well Party A wins, and party B loses so it deadlocks as simply party politics, because they are not trying to follow what people want, just change the system to benefit them.
IMO, the core issue is not the specifics it's a system which has been corrupted over time. Consider there is a North Dakota and South Dakota simply because that gives them more seats in the senate. So, now because of that power grab all those years ago we end up with a small population with more power than they would otherwise have and little reason to give it up.
Why does strict falsifiability not work in social sciences? You just need to state your theories better if you can otherwise find counterexamples easily. Physicists don't say "smash these two particles together and you'll see a Higgs Boson", they have sufficiently nuanced theories that sifting through petabytes of data to find what you claim exists is justified.
Take the Cornell Food Lab as example. One of the examples was that people will eat less sweets if they are stashed away and not fully visible on a table. It's interesting and it might very well be true, overall. But it is absurdly hard to know for sure. For one, devising an experiment for this is very hard. But the main problem is that you will absolutely find at least one person that eats more sweets when they are stashed away. So, the theory is wrong? Not really.
It might be an invalid theory though. But if you think that, there are very few things those fields could study. I'd say there were no use for them.
> they have sufficiently nuanced theories that sifting through petabytes of data to find what you claim exists is justified.
That does not sound like searching for falsifiability to me. However, it sounds like the workaround I also tried: Backing a theory - that expects an overall result, not to be true for each individual - up with as much data as possible so one can reasonable assume it is true. But that's not really the correct approach.
No, for something like the Higg Boson they try to see it or its effects (and if they were not to see it in the right circumstances, the theory were false). Might be a bad example though given how it touches physicists theory building.
Is a theory that is true only part of the time actually any use to anyone?
For example, your example. I might be the guy who eats more sweets when they're hidden. I go to my psych to talk about my sweet problem. The psych says "it's a well-studied phenomenon that you will eat less sweets if they're hidden". I hide my sweets. Boom, I eat more. What happens then?
We know there's a huge variety in humans. But statistics weeds them out and talks about "the average human". Which is great for statistics and researchers. But since there's no such thing as the average human, how does this actually help us?
If, say, 95% of people tend to eat less sweets when they're hidden, it's worth trying the method first (then discontinuing it if it you turn out to be one of the 5%). Is a drug that cures most, but not all, instances of a specific infection useless?
> Why does strict falsifiability not work in social sciences?
Theories in physics provide specific numerical predictions. Theories in social psychology provide predictions like 'X is correlated with Y' which is very difficult to falsify because in practice everything is correlated with everything else to some degree.
Null hypothesis significance testing is not the same as Popperian falsification either - it attempts to falsify the null hypothesis rather than the theory in question, and in the social sciences the null hypothesis is almost never strictly true.
Meehl was talking about this as far back as 1990:
> Null hypothesis testing of correlational predictions from weak substantive theories in soft psychology is subject to the influence of ten obfuscating factors whose effects are usually (1) sizeable, (2) opposed, (3) variable, and (4) unknown
The consensus among folk who care about the problem seems to be that we should move from significance testing to predicting and estimating effect sizes.
And then of course on top of that we still have to deal with publication bias, HARKing, researcher degrees of freedom etc.
I mean, of course the big problem is that applying strict falsifiability to social sciences does not work. You can always find a counter-example, the problems studied do not work like that. It is hard to reconcile this, but I think that is the heart of the replication crisis (together with some bad statistics). But (having done a PhD in a somewhat related field) I see only two options:
1. We need a complete overhaul of how experiments and study results are published, so that observers can see the failed results and we can try to assess how often a theory hold up
2. We have to limit non-hard sciences to questions that are not ambiguous, where one well-done negative result really shows a theory is wrong.
Of course, there is a third option: Just continue as it is now and ignore all results that seem unlikely, because given how those fields work, they are most likely wrong.