Is inter-project supervision possible?

All great points related to the technical capabilities of OTP supervisors, including points I haven’t internalized yet. So we agree that the technical objective of running cross node supervision trees exists; your detailed earlier post also made that clear.

My larger point was that just because you can doesn’t mean you should and knowing if you should or not often comes with having done a fair amount of research beyond just can it be made to work. You’re right in pointing out that understanding the failure modes is necessary to which I’d add understanding the “failure domains” of your application: the boundaries within which you can have failures which can be sufficiently isolated as to allow other application services to continue without interruption. But this leads me back to the original poster’s stated objective and architecture:

Having two devices, one device watching another for failure, is absolutely a thing and is appropriate in some scenarios… one could easily argue that the whole idea of supervisors is this concept simply abstracted into software and within an application. Where I think the original poster errs is that it appears they’re looking at a tool which accomplishes the fine-grained supervision of application processes, rather than protecting against the more general “is that other device still responsive or not” type failures. I don’t think you become more fault tolerant in those circumstances: you can’t achieve higher fault tolerance by more tightly coupling application processes across nodes which necessarily introduces new faulting scenarios which come with multiple devices and communication buses; @lud makes some excellent points about additional failure modes introduced by the proposed approach. If you do want “device a watches/manages device b”, I still think you’d be better off making the application services running on each node as independent as possible rather than deepening the coupling… application processes being supervised locally within the node… while creating/getting a library for a more fit-for-purpose service to allow device a and device b to communicate state and command restarts and such, avoiding the process linking which comes with OTP supervisors.

I’m happy to be shown that I’m wrong about my assumptions here, but my prior experiences make me weary of this kind of coupling.

In either case, to actually achieve higher and not lesser, fault tolerance, I still think there are payoffs in time and effort to doing some up-front research on the fundamentals. I think its reasonable for a newcomer to see supervisors and think maybe they can/should be used this way… but I think they’d need to next also be asking themselves the kinds of questions I originally posted before committing too far to that path. Yes, some of those questions will have satisfactory answers as you’ve pointed out, but others maybe not so much.