囍 is a really bad example of Unicode bloat since it's actually used very frequently.
The real issue are the thousands of characters that appeared incidentally in some ancient text, either as typo or as a weird interpretation of a common character, which then ended up in the Kangxi dictionary, and then subsequently imported en-masse into Unicode.
Example, 𠒇 - the only known use (in non-ancient times) of this character was being the official name of 雷莊𠒇, winner of the 2017 Miss Hong Kong Pageant. According to her, she intended to write 雷莊兒 when she applied for her official documents, but somehow the officials interpreted it as 𠒇, which is really an archaic form of 兒 (at best). Reportedly she's changing her name back to 雷莊兒
Thousands of such characters exist, if you look at the page where 𠒇 is supposed to originate, more than half of this is obsolete -- https://www.kangxizidian.com/v1/?page=125#gv
(So yeah, you're correct in essence but picking on 囍 as an example probably doesn't really get your point across...)
You've completely missed my point. Unicode is only intended to represent writing systems. It doesn't matter whether they're ancient, modern, common, or rare, but they're supposed to be writing systems.
囍 is used frequently, but it is not part of any writing system and does not convey any linguistic message. The opposite is true for 𠒇 - it is not used frequently, but it is part of a writing system and is used to convey linguistic messages.
The reason the point about 𠒇 holds is that it really isn't a character. I didn't claim it was "rare", I said it was a typo or a misinterpretation of an actual character (for the case of 𠒇 it is a misrepresented form of 兒.)
To presume a character is "real" merely because it exists in the Kangxi dictionary is as valid reasoning as presuming a character is "real" because it exists in "some other encoding standard" that you've been dismissive about. It's just that Kangxi is the de-facto Han character encoding scheme before computer encodings were invented. (Unihan even contains all details about the radical, stroke and even page number where it was sourced from)
A lot of those characters appeared once in some ancient text, and took the Kangxi form due to transcriptions from scribes across the centuries, but we actually have no evidence that they are "real" (at any point in time). Some of these characters are known alternative forms of common characters, or are only known to appear in some ancient text before Han characters were standardized. Some are plain typographical errors. It's like a 3 year old child learning to write "ABC", which looks a bit weird, and then the unicode committee assigned 3 code points to them.
> Some of these characters are known alternative forms of common characters, or are only known to appear in some ancient text before Han characters were standardized.
Those are entirely valid for Unicode. Han unification in Unicode is already considered a mistake. That's why "unified" code points now also have explicit, higher-numbered 'equivalent' code points that unambiguously refer to a particular graphical form. The graphical form is the whole point of Unicode.
The real issue are the thousands of characters that appeared incidentally in some ancient text, either as typo or as a weird interpretation of a common character, which then ended up in the Kangxi dictionary, and then subsequently imported en-masse into Unicode.
Example, 𠒇 - the only known use (in non-ancient times) of this character was being the official name of 雷莊𠒇, winner of the 2017 Miss Hong Kong Pageant. According to her, she intended to write 雷莊兒 when she applied for her official documents, but somehow the officials interpreted it as 𠒇, which is really an archaic form of 兒 (at best). Reportedly she's changing her name back to 雷莊兒
Thousands of such characters exist, if you look at the page where 𠒇 is supposed to originate, more than half of this is obsolete -- https://www.kangxizidian.com/v1/?page=125#gv
(So yeah, you're correct in essence but picking on 囍 as an example probably doesn't really get your point across...)