The large majority of what you've written here is valid, but I think you're wrong in the conclusion that the model is never (or almost never) useful.
I think this disagreement is going to come down to two things:
[Owen] thinks for problems where success is made by a bunch of small contributions this is theoretically justified. A good example of such a field is research. And studies have indeed shown that research returns are indeed logarithmic. It just doesn't apply to *every* field
I wonder if the disclosures could be non-text by default -- e.g. colour-coded with an optional footnote for details.
The thing I'm not liking as a reader is having words to process on this stuff at the start (for me this isn't just cases where people aren't following policy; I've felt it some about a case where the words were one of the suggested wordings from the policy). Non-text ways to signal could potentially get best-of-both-worlds in terms of reader attention.
Ok so I can kind of tune into what you're saying here, but I also feel kind of uneasy about it. I guess I'd be curious what you make of the following potential arguments:
Ok, so one place the predictions of these theories might come apart is that my theory suggests a norm against impersonating medics, whereas I think yours doesn't (although maybe I'm just not seeing it; I don't think I would have said that avoiding torture of prisoners was part of protecting the mechanisms of ending war, although I do kind of see what you mean). I haven't looked into it at all, but if that norm has emerged independently multiple times that would be suggestive in favour of the broader theory; whereas if it has just emerged once it looks perhaps more potentially-idiosyncratic, which would be suggestive in favour of the narrower theory.
I agree that the model I proposed is imprecise; I think this counts against its usefulness but not its validity.
I'm not suggesting this as a thing to advocate for; merely as a descriptive pattern of what the category of war crimes is doing. I think the things which make ending war harder are an important class of really destructive thing, but it seems clarity-obscuring to me to claim that this is definitionally what war crimes are? Rather than giving your thing a new label and then getting to discuss what fraction of war crimes are in that category, and whether there are things in that category which aren't war crimes (e.g. if torturing POWs counts under your categorization, then why doesn't conscription count -- after all, it damages the "one side runs out of soldiers" mechanism for ending war).
I like the puzzle. But I wonder if you can make your answer even simpler:
I think this explains the category that you outline (undermining trust in the kind of institutions that could stop the war is super destructive!), but also explains some other cases, e.g. abuse of prisoners, not impersonating medical staff, etc.
Hmm, I've used LLMs to varying degrees in writing articles. Usually not to the point of writing significant amounts of text, but a case where I think it clearly helped to improve the output is this story: https://strangecities.substack.com/p/some-days-soon
Functionally, I wrote a complete draft, then got Claude to redraft, then I went through and stitched the best bits of the two drafts together (or wrote new versions where that seemed best). (If you thought the original draft was better I'd be interested to hear that: https://docs.google.com/document/d/1icY2wpcgvKszfzHFButKcOwV8B9xMypTAk48kjnOGz0/edit?usp=drivesdk )
(I notice that I'm more likely to find LLMs helpful in drafting things when writing fiction. I think it's least likely to help when it's important to convey my precise epistemic status towards the things I'm saying.)
Requiring disclosures to be at the top of the post (rather than e.g. allowing them to be at the bottom) does feel like it's sending some implicit "this is kind of bad so people need to be warned about it" message, even if it's in a "recommended uses" section.
Like I think people might reasonably worry about others pre-judging posts with this disclaimer, and hence (perhaps, sometimes) prefer workflows where they don't need to include the disclaimer, even if this makes their posts worse.
I don't think there's an easy answer here -- like, presumably the point of the policy is to allow this kind of pre-judging and let people make differently-informed choices about what they engage with. But I think the post kind of papers over this tension.
Re. 1, I agree that you're going to have the standard difficulties, but I think that the framework makes it easier to make somewhat-informed guesses (at least if you're happy to let us assume that returns are approximately logarithmic).
For instance, suppose you're wondering about contributing $100k to a field that currently spends $50M/year. What will that $100k buy? It's kind of hard to project counterfactuals. Well, what about doubling the size of the whole field? Or 1000x-ing the size of the field? Sometimes it can be easier to find a scale where effects feel like now you can bring different comparisons or intuitions to bear, as a way of triangulating on a reasonable number.
Re. 2, one point is that if we take log returns as given, then the tractability term for a given field will be the same, regardless of current investment. That feels like a helpful fact for reasoning about things? e.g. if you're comparing two research fields, you might insert a vibes-based judgement about their relative difficulty (especially if they're not wildly different in seeming difficulty), without worrying about if that's somehow getting swamped by differences due to the different sizes of the fields. Similarly it lets you make arguments like most problems fall within a 100x tractability range, which is sometimes enough to feel like it's doing something useful.
Now I don't think that log returns is always the right model (especially in domains where we have a good understanding of the interventions, and we're mostly concerned with their direct effects). But maybe that's enough to give you some of the flavour of why I think it can get you something helpful, if you do assume it? (I also think there are lots of ways to go wrong within the ITN framework, and endorse Lizka's post that I linked in my last comment.)