Hard to Vary or Hardly Usable?
#3069·Dennis HackethalOP revised 10 months agoMy critique of David Deutsch’s The Beginning of Infinity as a programmer. In short, his ‘hard to vary’ criterion at the core of his epistemology is fatally underspecified and impossible to apply.
Deutsch says that one should adopt explanations based on how hard they are to change without impacting their ability to explain what they claim to explain. The hardest-to-change explanation is the best and should be adopted. But he doesn’t say how to figure out which is hardest to change.
A decision-making method is a computational task. He says you haven’t understood a computational task if you can’t program it. He can’t program the steps for finding out how ‘hard to vary’ an explanation is, if only because those steps are underspecified. There are too many open questions.
So by his own yardstick, he hasn’t understood his epistemology.
You will find that and many more criticisms here: https://blog.dennishackethal.com/posts/hard-to-vary-or-hardly-usable
Brett says “Rational decision making is not a matter of pulling a lever and cranking through a calculation.”
#5415·Dennis HackethalOP, 27 days agoBrett says the bounty is “impossible” because “it’s asking for a definition for something that *cannot be defined* in the way the challenge demands.”
Dirk replies that the bounty doesn’t ask for essentialist definitions. It instead asks for clarity around how to actually use HTV.
#3069·Dennis HackethalOP revised 10 months agoMy critique of David Deutsch’s The Beginning of Infinity as a programmer. In short, his ‘hard to vary’ criterion at the core of his epistemology is fatally underspecified and impossible to apply.
Deutsch says that one should adopt explanations based on how hard they are to change without impacting their ability to explain what they claim to explain. The hardest-to-change explanation is the best and should be adopted. But he doesn’t say how to figure out which is hardest to change.
A decision-making method is a computational task. He says you haven’t understood a computational task if you can’t program it. He can’t program the steps for finding out how ‘hard to vary’ an explanation is, if only because those steps are underspecified. There are too many open questions.
So by his own yardstick, he hasn’t understood his epistemology.
You will find that and many more criticisms here: https://blog.dennishackethal.com/posts/hard-to-vary-or-hardly-usable
#5325·Dennis HackethalOP, about 1 month agoBrett Hall responds to the contradiction between HTV and ‘if you can’t program it, you haven’t understood it’: https://x.com/ToKTeacher/status/2090780593739243793
He agrees we can’t program HTV but says even if we don’t know how or why an explanation is HTV, we still know that it’s HTV.
The blog post already sidesteps that issue by granting the user the ability to input some HTV score without having to give any reasoning. That still leads to all sorts of trouble.
#5325·Dennis HackethalOP, about 1 month agoBrett Hall responds to the contradiction between HTV and ‘if you can’t program it, you haven’t understood it’: https://x.com/ToKTeacher/status/2090780593739243793
He agrees we can’t program HTV but says even if we don’t know how or why an explanation is HTV, we still know that it’s HTV.
Translation: ‘trust me bro’
It’s just vibes.
#3069·Dennis HackethalOP revised 10 months agoMy critique of David Deutsch’s The Beginning of Infinity as a programmer. In short, his ‘hard to vary’ criterion at the core of his epistemology is fatally underspecified and impossible to apply.
Deutsch says that one should adopt explanations based on how hard they are to change without impacting their ability to explain what they claim to explain. The hardest-to-change explanation is the best and should be adopted. But he doesn’t say how to figure out which is hardest to change.
A decision-making method is a computational task. He says you haven’t understood a computational task if you can’t program it. He can’t program the steps for finding out how ‘hard to vary’ an explanation is, if only because those steps are underspecified. There are too many open questions.
So by his own yardstick, he hasn’t understood his epistemology.
You will find that and many more criticisms here: https://blog.dennishackethal.com/posts/hard-to-vary-or-hardly-usable
Brett Hall responds to the contradiction between HTV and ‘if you can’t program it, you haven’t understood it’: https://x.com/ToKTeacher/status/2090780593739243793
He agrees we can’t program HTV but says even if we don’t know how or why an explanation is HTV, we still know that it’s HTV.
#5183·Dennis HackethalOP, about 2 months ago@edwin-de-wit recently asked Deutsch about this:
… hard to vary can also be a property that you don’t want. Sometimes they say that complicated systems are just so entangled that they’re no longer changeable. You could call that hard to vary but it’s not a property you pursue.
Deutsch’s reply:
Quite so. So hard to vary is about explanations, not about end products. So the end product should be easy to vary, if you want to. But it should be hard to vary in the sense that you don’t want to.
Naval, who’s inspired by Deutsch, says good products are hard to vary, too: https://nav.al/good-products
Gives iPhone as example.
#5183·Dennis HackethalOP, about 2 months ago@edwin-de-wit recently asked Deutsch about this:
… hard to vary can also be a property that you don’t want. Sometimes they say that complicated systems are just so entangled that they’re no longer changeable. You could call that hard to vary but it’s not a property you pursue.
Deutsch’s reply:
Quite so. So hard to vary is about explanations, not about end products. So the end product should be easy to vary, if you want to. But it should be hard to vary in the sense that you don’t want to.
But HTV is also about art (BoI chapter 14).
A piece of art is a kind of end product. And we could write code to produce art.
#5183·Dennis HackethalOP, about 2 months ago@edwin-de-wit recently asked Deutsch about this:
… hard to vary can also be a property that you don’t want. Sometimes they say that complicated systems are just so entangled that they’re no longer changeable. You could call that hard to vary but it’s not a property you pursue.
Deutsch’s reply:
Quite so. So hard to vary is about explanations, not about end products. So the end product should be easy to vary, if you want to. But it should be hard to vary in the sense that you don’t want to.
Explanations are functions. In other words, explanations are software. There’s no sharp distinction between explanations and end products here. So the concern about tight coupling still applies.
#3718·Dennis HackethalOP, 9 months agoFrom my article:
[D]epending on context, being hard to change can be a bad thing. For example, ‘tight coupling’ is a reason software can be hard to change, and it’s considered bad because it reduces maintainability.
@edwin-de-wit recently asked Deutsch about this:
… hard to vary can also be a property that you don’t want. Sometimes they say that complicated systems are just so entangled that they’re no longer changeable. You could call that hard to vary but it’s not a property you pursue.
Deutsch’s reply:
Quite so. So hard to vary is about explanations, not about end products. So the end product should be easy to vary, if you want to. But it should be hard to vary in the sense that you don’t want to.
Paul Raymond-Robichaud tells me (I think in this space) that in math, you can have a formalized notion of HTV. From what I recall he said, it sounded like that would be trivial to develop (for a mathematician like Paul – not for me).
Anyone who’s both mathematically and epistemologically inclined could give this a shot at generalization.
Space where Tyler, Charlie, and I discuss hard to vary: https://x.com/dchackethal/status/2083018149151084792
This could be a promising approach to formalize HTV:
https://x.com/FZdyb/status/2051352500582641931
https://github.com/deoxyribose/hard_to_vary_posterior_predictive
It’s AI generated, so not eligible for the bounty. And I’m not familiar enough with the probability calculus to evaluate it. But bookmarking it here for the future.
(One of my first criticisms was that HTV has nothing to do with likelihood, which the author granted but addressed by saying it maps onto marginal likelihood. See https://x.com/FZdyb/status/2051004605601898561 and surrounding discussion.)
We can redefine ‘hard to vary’, but we’d need still a working implementation in the form of computer code.
… Demeter scores 25% and axial tilt scores 100%.
Now do this universally, for any given theory.
We can redefine ‘hard to vary’, but we’d need still a working implementation in the form of computer code.
… Demeter scores 25% and axial tilt scores 100%.
Now do this universally, for any given theory.
#4940·Dennis HackethalOP, 5 months agoWe can redefine ‘hard to vary’, but we’d need still a working implementation in the form of computer code.
… Demeter scores 25% and axial tilt scores 100%.
Now do this universally, for any given theory.
By the way Knut, when I go straight into ‘criticism mode’, that can sound cold or harsh. But don’t let that discourage you from exploring your idea further. Maybe you’re onto something! A working implementation of hard to vary would be useful and vindicating.
#4936·Knut Sondre Sæbø revised 5 months agoHave some thoughts, which might be way off. But interested in your response. It seems to me that "hard-to-vary" is itself the criterion that a theory should be as programmable as possible. As you note, the goal of a theory should be to make it as explicit as possible, and a program is explicitness in its most complete form. Any theory with ambiguous components automatically has a breaking point that is changeable which is hard to detect. A programmable theory has strict causal relations all the way from the axioms to the prediction, which makes any change to the components detectable. In other words: a theory is hard to vary to the extent that its components and the couplings between them can be specified as a program. If a theory is vague, you cannot tell when it has been varied.
This might give a concrete operationalization. A breaking point is any place in the formalization where the chain stops being programmable: a primitive with no implementable type, a coupling between components that cannot be turned into a function, or just a step that requires implicit theories to fill the explanatory gaps. A mathematical theory with no remaining gaps has zero breaking points and is maximally hard to vary. A theory in natural language is already worse, because words carry ambiguity and vary from mind to mind. This does not rule out better and worse theories in natural language, since we can use more or less ambiguous words and relations. But it does create a hierarchy of hard-to-vary explanations, where the share of the explanation that is programmable, or at least unambiguous, forms the basis for measuring the "hard-to-vary" criterion.
This is probably too crude a formalization. But evaluating the two theories of Demeter's emotions and axial tilt as explanations, you could check how much of each is programmable. Detecting seasons is programmable in both cases through temperature and changes in weather. Demeter's emotions and the causal link from them to the weather, which is the entire explanation, are not programmable. In the axial tilt theory, every component is. So on this measure Demeter scores 25% and axial tilt scores 100%.
We can redefine ‘hard to vary’, but we’d need still a working implementation in the form of computer code.
… Demeter scores 25% and axial tilt scores 100%.
Now do this universally, for any given theory.
Have some thoughts, which might be way off. But interested in your response. It seems to me that "hard-to-vary" is itself the criterion that a theory should be as programmable as possible. As you note, the goal of a theory should be to make it as explicit as possible, and a program is explicitness in its most complete form. Any theory with ambiguous components automatically has a breaking point that is changeable without detection. A programmable theory has strict causal relations all the way from the axioms to the prediction, which makes any change to the components detectable. In other words: a theory is hard to vary to the extent that its components and the couplings between them can be specified as a program. If a theory is vague, you cannot tell when it has been varied.
This gives a concrete operationalization. A breaking point is any place in the formalization where the chain stops being programmable: a primitive with no implementable type, a coupling between components that cannot be specified, or a step that requires implicit theories to fill the explanatory gaps. A mathematical theory with no remaining gaps has zero breaking points and is maximally hard to vary. A theory in natural language is already worse, because words carry ambiguity and vary from mind to mind. This does not rule out better and worse theories in natural language, since we can use more or less ambiguous words and relations. But it does create a hierarchy of hard-to-vary explanations, where the share of the explanation that is programmable, or at least unambiguous, forms the basis for the criterion.
This is probably too crude a formalization. But evaluating the two theories of Demeter's emotions and axial tilt as explanations, you could check how much of each is programmable. Detecting seasons is programmable in both cases through temperature and changes in weather. Demeter's emotions and the causal link from them to the weather, which is the entire explanation, are not programmable. In the axial tilt theory, every component is. So on this measure Demeter scores 25% and axial tilt scores 100%.
Have some thoughts, which might be way off. But interested in your response. It seems to me that "hard-to-vary" is itself the criterion that a theory should be as programmable as possible. As you note, the goal of a theory should be to make it as explicit as possible, and a program is explicitness in its most complete form. Any theory with ambiguous components automatically has a breaking point that is changeable which is hard to detect. A programmable theory has strict causal relations all the way from the axioms to the prediction, which makes any change to the components detectable. In other words: a theory is hard to vary to the extent that its components and the couplings between them can be specified as a program. If a theory is vague, you cannot tell when it has been varied.
This might give a concrete operationalization. A breaking point is any place in the formalization where the chain stops being programmable: a primitive with no implementable type, a coupling between components that cannot be turned into a function, or just a step that requires implicit theories to fill the explanatory gaps. A mathematical theory with no remaining gaps has zero breaking points and is maximally hard to vary. A theory in natural language is already worse, because words carry ambiguity and vary from mind to mind. This does not rule out better and worse theories in natural language, since we can use more or less ambiguous words and relations. But it does create a hierarchy of hard-to-vary explanations, where the share of the explanation that is programmable, or at least unambiguous, forms the basis for measuring the "hard-to-vary" criterion.
This is probably too crude a formalization. But evaluating the two theories of Demeter's emotions and axial tilt as explanations, you could check how much of each is programmable. Detecting seasons is programmable in both cases through temperature and changes in weather. Demeter's emotions and the causal link from them to the weather, which is the entire explanation, are not programmable. In the axial tilt theory, every component is. So on this measure Demeter scores 25% and axial tilt scores 100%.
It seems to me that "hard-to-vary" is itself the criterion that a theory should be as programmable as possible. As you note, the goal of a theory should be to make it as explicit as possible, and a program is explicitness in its most complete form. Any theory with ambiguous components automatically has a breaking point that is changeable without detection. A programmable theory has strict causal relations all the way from the axioms to the prediction, which makes any change to the components detectable. In other words: a theory is hard to vary to the extent that its components and the couplings between them can be specified as a program. If a theory is vague, you cannot tell when it has been varied.
This gives a concrete operationalization. A breaking point is any place in the formalization where the chain stops being programmable: a primitive with no implementable type, a coupling between components that cannot be specified, or a step that requires implicit theories to fill the explanatory gaps. A mathematical theory with no remaining gaps has zero breaking points and is maximally hard to vary. A theory in natural language is already worse, because words carry ambiguity and vary from mind to mind. This does not rule out better and worse theories in natural language, since we can use more or less ambiguous words and relations. But it does create a hierarchy of hard-to-vary explanations, where the share of the explanation that is programmable, or at least unambiguous, forms the basis for the criterion.
This is probably too crude a formalization. But evaluating the two theories of Demeter's emotions and axial tilt as explanations, you could check how much of each is programmable. Detecting seasons is programmable in both cases through temperature and changes in weather. Demeter's emotions and the causal link from them to the weather, which is the entire explanation, are not programmable. In the axial tilt theory, every component is. So on this measure Demeter scores 25% and axial tilt scores 100%.
Have some thoughts, which might be way off. But interested in your response. It seems to me that "hard-to-vary" is itself the criterion that a theory should be as programmable as possible. As you note, the goal of a theory should be to make it as explicit as possible, and a program is explicitness in its most complete form. Any theory with ambiguous components automatically has a breaking point that is changeable without detection. A programmable theory has strict causal relations all the way from the axioms to the prediction, which makes any change to the components detectable. In other words: a theory is hard to vary to the extent that its components and the couplings between them can be specified as a program. If a theory is vague, you cannot tell when it has been varied.
This gives a concrete operationalization. A breaking point is any place in the formalization where the chain stops being programmable: a primitive with no implementable type, a coupling between components that cannot be specified, or a step that requires implicit theories to fill the explanatory gaps. A mathematical theory with no remaining gaps has zero breaking points and is maximally hard to vary. A theory in natural language is already worse, because words carry ambiguity and vary from mind to mind. This does not rule out better and worse theories in natural language, since we can use more or less ambiguous words and relations. But it does create a hierarchy of hard-to-vary explanations, where the share of the explanation that is programmable, or at least unambiguous, forms the basis for the criterion.
This is probably too crude a formalization. But evaluating the two theories of Demeter's emotions and axial tilt as explanations, you could check how much of each is programmable. Detecting seasons is programmable in both cases through temperature and changes in weather. Demeter's emotions and the causal link from them to the weather, which is the entire explanation, are not programmable. In the axial tilt theory, every component is. So on this measure Demeter scores 25% and axial tilt scores 100%.
I'm not a programmer, so the code below is 100% AI-generated. But here is an attempt. If we normalize a theory into the parts that can be put on a computer, the types it uses, the nodes (specific values) it commits to, and the functions between them, we can score the theory by how many of those parts run.
from dataclasses import dataclass
@dataclass
class Item:
name: str
kind: str # "type", "node", or "function"
programmable: bool # does it compile and run?
@dataclass
class Theory:
name: str
items: list[Item]
def score(self) -> float:if not self.items:return 0.0ok = sum(1 for i in self.items if i.programmable)return ok / len(self.items)
def compare(a: Theory, b: Theory) -> None:
print(f"{a.name:<20} {a.score():.0%}")
print(f"{b.name:<20} {b.score():.0%}")
Demeter theory
demeter = Theory("Demeter", [
Item("Latitude", "type", True),
Item("Temperature", "type", True),
Item("Goddess", "type", False),
Item("Emotion", "type", False),
Item("Demeter", "node", False),
Item("emotionstate", "node", False),
Item("emotionat", "function", False),
Item("emotion_weather", "function", False),
])
Axial tilt theory
axialtilt = Theory("Axial tilt", [
Item("Latitude", "type", True),
Item("Temperature", "type", True),
Item("Angle", "type", True),
Item("Insolation", "type", True),
Item("axialtilt", "node", True),
Item("solarconstant", "node", True),
Item("solarangle", "function", True),
Item("insolation_at", "function", True),
Item("temperature", "function", True),
])
compare(demeter, axial_tilt)
Output metrics:
Demeter 25%
Axial tilt 100%
If we normalize a theory into the parts that can be put on a computer, the types it uses, the nodes (specific values) it commits to, and the functions between them, we can score the theory by how many of those parts run.
The program goes through each item in a theory, counts how many are marked True, and divides by the total. That fraction is the score of how hard the theory is to vary.
An item is True if it can be put on a computer, either by reusing an existing type (Float, Int) or by defining a new one that compiles. It is False if no working type system can express it. The user fills in the labels; the program just counts.
Demeter: 2 of 8 items program (Latitude and Temperature). Demeter, her emotions, and the functions linking them to weather can't. So the score is 25%.
Axial tilt: 9 of 9 items program. Standard types, measured constants, and functions from standard physics. So the score is 100%.
If we normalize a theory into the parts that can be put on a computer, the types it uses, the nodes (specific values) it commits to, and the functions between them, we can score the theory by how many of those parts actually run.
from dataclasses import dataclass
@dataclass
class Item:
name: str
kind: str # "type", "node", or "function"
programmable: bool # does it compile and run?
@dataclass
class Theory:
name: str
items: list[Item]
def score(self) -> float:if not self.items:return 0.0ok = sum(1 for i in self.items if i.programmable)return ok / len(self.items)
def compare(a: Theory, b: Theory) -> None:
print(f"{a.name:<20} {a.score():.0%}")
print(f"{b.name:<20} {b.score():.0%}")
Demeter theory
demeter = Theory("Demeter", [
Item("Latitude", "type", True),
Item("Temperature", "type", True),
Item("Goddess", "type", False),
Item("Emotion", "type", False),
Item("Demeter", "node", False),
Item("emotionstate", "node", False),
Item("emotionat", "function", False),
Item("emotion_weather", "function", False),
])
Axial tilt theory
axialtilt = Theory("Axial tilt", [
Item("Latitude", "type", True),
Item("Temperature", "type", True),
Item("Angle", "type", True),
Item("Insolation", "type", True),
Item("axialtilt", "node", True),
Item("solarconstant", "node", True),
Item("solarangle", "function", True),
Item("insolation_at", "function", True),
Item("temperature", "function", True),
])
compare(demeter, axial_tilt)
Output metrics:
Demeter 25%
Axial tilt 100%
I'm not a programmer, so the code below is 100% AI-generated. But here is an attempt. If we normalize a theory into the parts that can be put on a computer, the types it uses, the nodes (specific values) it commits to, and the functions between them, we can score the theory by how many of those parts run.
from dataclasses import dataclass
@dataclass
class Item:
name: str
kind: str # "type", "node", or "function"
programmable: bool # does it compile and run?
@dataclass
class Theory:
name: str
items: list[Item]
def score(self) -> float:if not self.items:return 0.0ok = sum(1 for i in self.items if i.programmable)return ok / len(self.items)
def compare(a: Theory, b: Theory) -> None:
print(f"{a.name:<20} {a.score():.0%}")
print(f"{b.name:<20} {b.score():.0%}")
Demeter theory
demeter = Theory("Demeter", [
Item("Latitude", "type", True),
Item("Temperature", "type", True),
Item("Goddess", "type", False),
Item("Emotion", "type", False),
Item("Demeter", "node", False),
Item("emotionstate", "node", False),
Item("emotionat", "function", False),
Item("emotion_weather", "function", False),
])
Axial tilt theory
axialtilt = Theory("Axial tilt", [
Item("Latitude", "type", True),
Item("Temperature", "type", True),
Item("Angle", "type", True),
Item("Insolation", "type", True),
Item("axialtilt", "node", True),
Item("solarconstant", "node", True),
Item("solarangle", "function", True),
Item("insolation_at", "function", True),
Item("temperature", "function", True),
])
compare(demeter, axial_tilt)
Output metrics:
Demeter 25%
Axial tilt 100%
#4931·Knut Sondre Sæbø, 5 months agoIt seems to me that "hard-to-vary" is itself the criterion that a theory should be as programmable as possible. As you note, the goal of a theory should be to make it as explicit as possible, and a program is explicitness in its most complete form. Any theory with ambiguous components automatically has a breaking point that is changeable without detection. A programmable theory has strict causal relations all the way from the axioms to the prediction, which makes any change to the components detectable. In other words: a theory is hard to vary to the extent that its components and the couplings between them can be specified as a program. If a theory is vague, you cannot tell when it has been varied.
This gives a concrete operationalization. A breaking point is any place in the formalization where the chain stops being programmable: a primitive with no implementable type, a coupling between components that cannot be specified, or a step that requires implicit theories to fill the explanatory gaps. A mathematical theory with no remaining gaps has zero breaking points and is maximally hard to vary. A theory in natural language is already worse, because words carry ambiguity and vary from mind to mind. This does not rule out better and worse theories in natural language, since we can use more or less ambiguous words and relations. But it does create a hierarchy of hard-to-vary explanations, where the share of the explanation that is programmable, or at least unambiguous, forms the basis for the criterion.
This is probably too crude a formalization. But evaluating the two theories of Demeter's emotions and axial tilt as explanations, you could check how much of each is programmable. Detecting seasons is programmable in both cases through temperature and changes in weather. Demeter's emotions and the causal link from them to the weather, which is the entire explanation, are not programmable. In the axial tilt theory, every component is. So on this measure Demeter scores 25% and axial tilt scores 100%.
If we normalize a theory into the parts that can be put on a computer, the types it uses, the nodes (specific values) it commits to, and the functions between them, we can score the theory by how many of those parts actually run.
from dataclasses import dataclass
@dataclass
class Item:
name: str
kind: str # "type", "node", or "function"
programmable: bool # does it compile and run?
@dataclass
class Theory:
name: str
items: list[Item]
def score(self) -> float:if not self.items:return 0.0ok = sum(1 for i in self.items if i.programmable)return ok / len(self.items)
def compare(a: Theory, b: Theory) -> None:
print(f"{a.name:<20} {a.score():.0%}")
print(f"{b.name:<20} {b.score():.0%}")
Demeter theory
demeter = Theory("Demeter", [
Item("Latitude", "type", True),
Item("Temperature", "type", True),
Item("Goddess", "type", False),
Item("Emotion", "type", False),
Item("Demeter", "node", False),
Item("emotionstate", "node", False),
Item("emotionat", "function", False),
Item("emotion_weather", "function", False),
])
Axial tilt theory
axialtilt = Theory("Axial tilt", [
Item("Latitude", "type", True),
Item("Temperature", "type", True),
Item("Angle", "type", True),
Item("Insolation", "type", True),
Item("axialtilt", "node", True),
Item("solarconstant", "node", True),
Item("solarangle", "function", True),
Item("insolation_at", "function", True),
Item("temperature", "function", True),
])
compare(demeter, axial_tilt)
Output metrics:
Demeter 25%
Axial tilt 100%
#3069·Dennis HackethalOP revised 10 months agoMy critique of David Deutsch’s The Beginning of Infinity as a programmer. In short, his ‘hard to vary’ criterion at the core of his epistemology is fatally underspecified and impossible to apply.
Deutsch says that one should adopt explanations based on how hard they are to change without impacting their ability to explain what they claim to explain. The hardest-to-change explanation is the best and should be adopted. But he doesn’t say how to figure out which is hardest to change.
A decision-making method is a computational task. He says you haven’t understood a computational task if you can’t program it. He can’t program the steps for finding out how ‘hard to vary’ an explanation is, if only because those steps are underspecified. There are too many open questions.
So by his own yardstick, he hasn’t understood his epistemology.
You will find that and many more criticisms here: https://blog.dennishackethal.com/posts/hard-to-vary-or-hardly-usable
It seems to me that "hard-to-vary" is itself the criterion that a theory should be as programmable as possible. As you note, the goal of a theory should be to make it as explicit as possible, and a program is explicitness in its most complete form. Any theory with ambiguous components automatically has a breaking point that is changeable without detection. A programmable theory has strict causal relations all the way from the axioms to the prediction, which makes any change to the components detectable. In other words: a theory is hard to vary to the extent that its components and the couplings between them can be specified as a program. If a theory is vague, you cannot tell when it has been varied.
This gives a concrete operationalization. A breaking point is any place in the formalization where the chain stops being programmable: a primitive with no implementable type, a coupling between components that cannot be specified, or a step that requires implicit theories to fill the explanatory gaps. A mathematical theory with no remaining gaps has zero breaking points and is maximally hard to vary. A theory in natural language is already worse, because words carry ambiguity and vary from mind to mind. This does not rule out better and worse theories in natural language, since we can use more or less ambiguous words and relations. But it does create a hierarchy of hard-to-vary explanations, where the share of the explanation that is programmable, or at least unambiguous, forms the basis for the criterion.
This is probably too crude a formalization. But evaluating the two theories of Demeter's emotions and axial tilt as explanations, you could check how much of each is programmable. Detecting seasons is programmable in both cases through temperature and changes in weather. Demeter's emotions and the causal link from them to the weather, which is the entire explanation, are not programmable. In the axial tilt theory, every component is. So on this measure Demeter scores 25% and axial tilt scores 100%.