Gelman, Bayesian Workflow:
A sometimes annoying habit of Bayesians is to take non-Bayesian methods and give them Bayesian interpretations. For example, “maximum likelihood is just Bayesian inference with a flat prior,” “fixed effects are just random effects with the group-level variance set to infinity,” “lasso is just regression with an exponential prior,” or “confidence intervals are just posterior intervals assuming flat priors.”
This practice can be useful in giving insight into the scenarios in which a statistical procedure will apply, but it can also be misleading. An example of the latter is the interpretation of lasso (Tibshirani 1996) mentioned just above: lasso is just regression with an exponential prior. Lasso is an approach for obtaining more stable regression estimates that pulls coefficient estimates toward or all the way to zero. The lasso estimate can be viewed as an approximate posterior mode assuming independent exponential priors on the coefficients, and we see some value in making this connection—but this does not make lasso a Bayesian method. The problem is that lasso is intended for use in problems of moderate and high dimensions. In such problems, the mode is not necessarily or even usually a good summary of the posterior distribution, because of the phenomenon of concentration of measure. As such, in high dimensional problems, the formulation, “a penalty is just a log prior density,” stops being useful or accurate. This does not mean that lasso is a bad idea, or that a Bayesian version is necessarily better; some of lasso’s desirable computational and applied properties directly arise from it being a mode rather than a full posterior.
More generally, any statistical method can be evaluated on its own, in non-Bayesian terms, by studying its statistical properties under various assumptions. Yes, least squares regression can be derived as the maximum likelihood estimate assuming independent normally-distributed errors with equal variance. But you can use least squares without making those assumptions, and you can evaluate the robustness of least squares under various alternative assumptions. That said, the theoretical derivation of least squares can provide some insight and practical guidance. For example, if you have reason to believe that the model errors have some particular covariance structure, you could consider incorporating that into an improved estimate. Conversely, if least squares estimation does not seem to be performing well, you could go back and check its implicit assumptions.
Still, careful Bayesian interpretations offer potential benefits. In applied settings where estimates are problematically noisy, leading to replication failures (Button et al. 2013), we can get some perspective by recognizing that simple means, differences, and least squares regression correspond to Bayesian estimates assuming flat priors. When the posterior distribution includes implausible parameter values, it can make sense to add further information. This can take the shape of an informative prior. Another way to say this: When the estimate is too large to believe, this implies that inference can be improved by incorporating prior information.
Yudkowsky, Searching for Bayes-Structure:
In fact, any part of a cognitive process that contributes usefully to truth-finding must have at least a little Bayesian structure—must harmonize with Bayes, at some point or another—must partially conform with the Bayesian flow, however noisily—despite however many disguising bells and whistles—even if this Bayesian structure is only apparent in the context of surrounding processes. Or it couldn’t even help.
(...)
But perhaps it is not quite as exciting to see something that doesn’t look Bayesian on the surface, revealed as Bayes wearing a clever disguise, if: (a) you don’t unravel the mystery yourself, but read about someone else doing it (Newton had more fun than most students taking calculus), and (b) you don’t realize that searching for the hidden Bayes-structure is this huge, difficult, omnipresent quest, like searching for the Holy Grail.
It’s a different quest for each facet of cognition, but the Grail always turns out to be the same. It has to be the right Grail, though—and the entire Grail, without any parts missing—and so each time you have to go on the quest looking for a full answer whatever form it may take, rather than trying to artificially construct vaguely hand-waving Grailish arguments. Then you always find the same Holy Grail at the end.
(...)
This left me in a bit of a pickle when it came to trying to explain in advance where I was going. I know from experience that if I say, “Bayes is the secret of the universe,” some people may say “Yes! Bayes is the secret of the universe!”; and others will snort and say, “How narrow-minded you are; look at all these other ad-hoc but amazingly useful methods, like regularized linear regression, that I have in my toolbox.”
(...)
To see through the surface adhockery of a cognitive process, to the Bayesian structure underneath—to perceive the probability flows, and know how, not just know that, this cognition too is Bayesian—as it always is—as it always must be—to be able to sense the Force underlying all cognition—this, is the Bayes-Sight.
So, to Gelman, the “Bayes-structure” is only sometimes helpful, but sometimes misleading as to why the thing works and in what ways it does, but to Yudkowsky, it is the Holy Grail and the only reason why these methods work is because of the Bayesian explanation underneath (Neatly, they both mention the same example as demonstrating the opposite things).
In my experience as a statistician, I think Gelman has the right of it (it is certainly the case with LASSO; I think Castillo is essentially correct that the LASSO posterior is a “useless object” that does away with why people use it), but I wonder about some more orthodox Bayesian perspectives on this contrast?
This is a classical Lakatosian conflict. In his terms, Alice is doing something called ‘monster-barring’: Alice has an (implicit) ‘conjecture’, Bob produces (counter)examples which Alice bars by adding conditions on what counts as an example. Your defense is also in Lakatos: the counterexamples which at first are barred are eventually synthesized into the ‘proof-generated concept’ which is the definition Alice did not have when she started (or the conjecture is bled into vacuity, which is informative either way). A more productive form of discourse is lemma-incorporation (identify what the counterexample actually refutes and write it as a condition to be satisfied).
There is also an answer there for how one detects whether or not the goalposts have actually shifted or if there really is a refinement towards a correct domain of validity: one observes whether or not the successive barrings decrease content without adding anything back. Since in her case the conjecture is a negative one (there does not exist an X), the content must come from the evidential standard (have you excluded every achievable way to prove X, such that you can no longer prove anything); is the standard applicable globally or is there a reason why it is applied locally; does the exclusion criteria kill only the specific study that was brought up or does it produce a rule for looking for what studies might count somewhere, etc.
Since the question and part of the answer match very closely, I think it would be sensible for you to read on this perspective (not a hard philosopher to read, nor uncommon; the thing to look at is Proofs & Refutations, which is a dialogue on V—E + F = 2, easy to digest for anyone with the typical backgrounds here)