There is no evidence of any semantic relation between this character and 𬌼 (U+2C33C). Also, the shapes are not identical. No unification without additional evidence establishing a relationship.
As you can see both in the GĐNHV example above and in this image from KCHN, Vietnam uses both forms. Not sure it's a good idea to unify, even if there is semantic overlap
The meaning is "drain dry", which is similar to 滗 (U+6ED7), “xế” is a native word, so this a case where a variant of 滗 was borrowed for its meaning. Unification should be appropriate.
The shapes are too different to recognize as identical. If we are going to arbitrarily equate simplified components based on interchangeability, then we should apply that across the board, including 馬/马, 金/钅, etc.
Oppose Unification
[ Unresolved from v1.0 ]
The point is not that we can derive correspondences, we can similarly derive correspondences from 馬 to 马, 鳥 to 鸟, etc. But, if we are going to say that because we can derive correspondences between glyphs that on the surface look quite different, then we should start using stronger unification that includes traditional and simplified. I don't think people want that, so the same treatment should be applied to simplified forms in languages other than Chinese used in the PRC.
The original source reference for V2-7A3B is Vũ Văn Kính, "Tự Điển Chữ Nôm", p. 272, shown in the image below. As you can see, the phonetic is 迭 (điệt). So, the current shape of U+28540 is incorrect. Unification will be acceptable if we change the shape of U+28540 to VN-F2173.
Unihan data, the ORT Attributes predictor, and most other candidates in WS2024 give 8. It would be better to be consistent.
Total Stroke Count
[ Unresolved from v1.0 ]
Given the variations across geographies and font designs, and the fact that unification precludes most shape-based determination of attributes, CJKJRG / IRG originally chose to use the Kangxi values, the most common denominator in dictionaries used by the CJKV countries. This avoided a lot of fruitless debate. Kangxi is 9 strokes, but as you point out, that later changed. I'm fine with either 8 or 9, but we should be consistent moving forward and change the ORT tools to support our decision. Otherwise, maybe we should just stop using TS.
The IRG Attributes Predictor counts 巨 as 5 strokes, Unihan has 4. We should discuss and document the stroke count we are going to use and fix the ORT if we decide it's 4. Otherwise keep TC=18.
Swapping is fine. But maybe we should consider 196 鳥, since it means 'crow'. We chose 86 based on Kangxi and Unihan, but if there is no need to follow those, 'bird' is the best semantic.
The IDS proposed above seems confusing. 㓁 is a variant of rad. 122 and always appears above. If we merely want to reduce the # of strokes, U+5197 would be better since it can have the shape ⿱冖儿.
秩 is the phonetic and 刀 the semantic. I don't see how radical 93 is appropriate here. If anything the secondary radical, taken from 秩, should be 115 (禾)
The IRG needs to have an intelligble policy on assignment of radicals. We originally based it on semantic, then the policy seems to have switched to "most intuitive". 180 is "intuitive"; 108 is semantic. Which is it to be?
The attributes predictor tool gives 8 for the stroke count. https://hc.jsecs.org/irg/ws2021/app/attributes-predictor.php?ids=%E2%BF%B1亡目务&radical=109.0
Evidence from a printed version of the same Manyōshū poem in an edition edited by Tsuru Hisashi and Moriyama Takashi showing emendation of UTC-03243 found in the Nishi Honganji manuscript to 霺 based on Ōya and Kyoto University manuscripts.
"Giúp đọc Nôm và Hán Việt" is currently the only evidence we have. But based on the analysis given in that dictionary, "Hv tâm quải", which means that it's composed of the Hán Việt characters 忄 and 挂, the glyph should be ⿰忄 挂.
This is currently the only example we have of this character. The component ⿻沈丶 is thought to derive from a simplified form of 㴷 (đắm: shipwrecked, see TĐCNTD p. 341), where 耽 has been reduced to the form V+60779 shown below:
The more common form is V+607C5 in the above, also shown here from the same source as VN-F2002:
I agree that the phonetic doesn't make sense. We would expect 卯 (mão), making this equivalent to U+24D60. TĐCNT is currently the only source we have, but will try to find more.
The reading is shown in the transliteration that on the page that follows (in green). The note explains that the original text is in Vietnamese, so no translation is given. "phượng" is the legendary bird typically translated as "phoenix". It is more commonly written: 鳯.
There are more than 40 glyphs using the same design in the NomNaTong font. It would be a significant effort to change them all. We would need to better understand the rationale for this design before making such a change.
The evidence is contradictory. The glyph has 永, but the structural analysis shows "băng giá", which in chữ Hán is 氷這. The character means "frost", so it has been corrected to use "ice" instead of "eternal".
There are 21 Vietnamese characters with 叕 as an immediate constituent. The distribution of the stroke shape in question is about half and half. We will investigate the issues with normalization.
Nom Na Tong and other Nôm fonts, such as Han-Nom Minh and Han-Nom Kai use 𥝢 for most of the characters shown above and some others:
Chars with 𥝢 in Nom Na Tong: 棃犂黎瓈𥗍𨛫㰀嚟𠠍𤂱𤑬
The only exceptions we can find are U+853E and VN-F03D1.
䄪 is not necessarily standard. Other dictionaries show 𥝢 for U+68C3 and U+853E, as in the entries below from Taberd.
Changing the glyphs of U+853E and VN-F03D1 will be the least disruptive and conform to Vietnamese usage. We can do the horizontal extension after we change U+853E.
We will add 𥝢 as a normalization rule for 䄪.
Entries from Taberd showing the use of 𥝢
P. 699
P. 260
Glyph design
We have an update to this glyph for the change proposed above.
Unfortunately a large number of the characters with element 戎 (U+620E, nhung) are designed this was in the reference font. This includes 戎 itself. We'll have to consider how to migrate these.
The phonetic, "dan" argues for U+67EC. Here is another analysis (Vũ Văn Kính, "Tự điễn chứ Nôm" p. 225) showing that the traditional and simplified forms both contain U+67EC, read "lan", as phonetic.
The element on the right is a simplification of the characters 沒 / 没, read "một", through these steps 没 > 𠬛 > 𠬠 or 𱥺 > 𠬠. There are 2 basic forms, 𠬠 and 𰰝. This is documented in the character definition shown in the image below from TĐCNTD p. 802
Below is an example of VN-F0CBC from "Lục Vân Tiên" showing a form somewhat between 𠬠 and 𰰝
Historically, there are many examples of 𰰝, but the current trend is to standardize on 𠬠, as shown in this the "BẢNG CHỮ HÁN NÔM CHUẨN THƯỜNG DÙNG" http://www.hannom-rcv.org/NS/bchnctd%20300623.pdf
About the comment #8629. I assume that the suggestion is to change the 亅 in the top 可 to 丨. That's a reasonable suggestion, but there are at least 6 other V-source characters already encoded (𣘁 U+24819, etc.) that use the current design. Since that's a majority, the lesser impact solution would be to normalize the remaining characters 哥 U+54E5, 歌 U+6B4C, U+2BC04, and VN-F176E (in WS2021) to use 亅.
Both variants are found in Vietnamese, TĐCNDG entry shown below has U+79C3 秃. In the NomNaTong font there are 5 glyphs composed with U+79C3 秃 and 5 composed with U+79BF 禿. Of the characters with V-Source references, if we normalize to 禿, we would also want to change 𥟉 U+257C9 / V3-3531 and 𥟹 U+257F9 / V2-7F31. If we normalize to 秃, we would only change U+22B33 / VN-22B33
We need to discuss attributes for the abbreviated component 𫇦 U+2B1E6. For most characters that use this, the radical is 140 with TC = 6 strokes. It would be good to follow that convention, with rad. 151 as secondary. Following that scheme, even with RS=151 as primary, SC=6, TC = 13, and FS = 2.
ORT reports 3 for FS, but I believe that wrong since the traditional phonetic for 吞 is 天, the first stroke of which is 1. This does not appear to be structured with the variant 呑 (U+5451), whose first stroke is 3.
What is the justification for labeling this as similar to U+31FC3?
Other
Evidence # 3 for UTC-03292, which has 逃入清化 (he fled into Thanh Hoá), parallels the phrase 奔清⿱花一 above and suggests that this character is a variant of 化 (U+5316, read hoá). Thanh Hoá is more commonly written 清化.
Both characters, U+23813 and VN-F0423 mean "a type of bamboo". The major sources, BTCN, ĐTĐCN, GĐNHV, KCHN, and Takeuchi, all show the form VN-F0423, with 竹, appropriately, as the radical. Since VN-F0423 appears to be the correct form, if we were to unify these, Vietnam would request changing the representative glyph for U+23813 to be that of VN-F0423.
IRG Working Set 2024v2.0
Source: Lee COLLINS
Date: Generated on 2026-07-27
Unification
Showing 20 comments.
This should have been unified with 𰜶 (U+30736) and withdrawn
The meaning is "drain dry", which is similar to 滗 (U+6ED7), “xế” is a native word, so this a case where a variant of 滗 was borrowed for its meaning. Unification should be appropriate.
蕈 (U+8548). Same semantic, very similar shape
Agree with unification. 釆 (biện) is often used for the phonetic 'thai', 'hai', more properly written 采 (thái)
The original source reference for V2-7A3B is Vũ Văn Kính, "Tự Điển Chữ Nôm", p. 272, shown in the image below. As you can see, the phonetic is 迭 (điệt). So, the current shape of U+28540 is incorrect. Unification will be acceptable if we change the shape of U+28540 to VN-F2173.
Attributes
Showing 164 comments.
Evidence
Showing 24 comments.
The more common form is V+607C5 in the above, also shown here from the same source as VN-F2002:
An appropriate normalization would be V+607C5.
V4-407A is currently encoded as 𫢠 U+2B8A0. One solution would be to move V4-407A to WS2024:00144 and change kIRG_VSource for U+2B8A0 to VN-2B8A0.
Glyph Design & Normalization
Showing 23 comments.
Nom Na Tong and other Nôm fonts, such as Han-Nom Minh and Han-Nom Kai use 𥝢 for most of the characters shown above and some others:
Chars with 𥝢 in Nom Na Tong: 棃犂黎瓈𥗍𨛫㰀嚟𠠍𤂱𤑬
The only exceptions we can find are U+853E and VN-F03D1.
䄪 is not necessarily standard. Other dictionaries show 𥝢 for U+68C3 and U+853E, as in the entries below from Taberd.
Changing the glyphs of U+853E and VN-F03D1 will be the least disruptive and conform to Vietnamese usage. We can do the horizontal extension after we change U+853E.
We will add 𥝢 as a normalization rule for 䄪.
Entries from Taberd showing the use of 𥝢
P. 699
P. 260
Below is an example of VN-F0CBC from "Lục Vân Tiên" showing a form somewhat between 𠬠 and 𰰝
Historically, there are many examples of 𰰝, but the current trend is to standardize on 𠬠, as shown in this the "BẢNG CHỮ HÁN NÔM CHUẨN THƯỜNG DÙNG" http://www.hannom-rcv.org/NS/bchnctd%20300623.pdf
Other
Showing 16 comments.
Data for Unihan
Showing 22 comments.