Yes, picchio monitors a running llama-server on a timer and flags when the prefill/decode ratio goes CPU-shaped. I hot-swapped one mid-run. probe 4 caught ENGAGED -> NOT ENGAGED.
tbh, backups matter. but nobody would accept Word deleting your files when you cancel Office. somewhere along the way we stopped distinguishing backup from custody.
i think this is mixing two separate ideas.
MTP is the training-side piece. speculative decoding is the inference trick. DeepSeek V3 used MTP as an auxiliary loss. the 2022 Google paper is speculative decoding. now Google is combining them.
https://arxiv.org/abs/2404.19737
Oh... so MTP is not speculative decoding? The (T)oken (P)rediction made me think it was on the inference side. I shall read the paper.
Edit: Ok, I understand now. You are saying that MTP has two aspects. 1) The training (for the mini-models to generate tokens), and 2) The actual speculative decoding implementation on the inference side (which uses those trained mini-models).
Not being familiar with the term CUA, I looked it up and TIL something (https://en.wikipedia.org/wiki/IBM_Common_User_Access). Since the doc you linked is dated 1988 and CUA was a brand new thing circa OS/2 (at least in traditional IBM timescales), the apparent inconsistency might be down to the massive scale of IBM in those days and organizational propagation delay.
Another factor could be still-reverberating echoes of the likely political battles around something as broad and far-reaching as CUA. I can only imagine the quiet boardroom battles won and lost fighting over CUA between different factions across all of IBM's kingdoms, divisions and principalities.
fwiw, understanding was never really the goal. calling it "debt" assumes all of it needs to be repaid, but most of the time this just feels like brain OOM. imo the skill is knowing what actually needs to stay in your head.
It seems to me that assuming that debt needs to be repaid was moved to a footnote during the ZIRP years. I feel like there was a shift of the poles, the movement of a rift, you get the idea.
IMO this probably isn't just about latency. keeping people in voice gives them training data text never will. is that why they were fine going transceiver over sfu and mostly ignoring multi-party?