MCPcopy Create free account
hub / github.com/IBM/Project_CodeNet / remove_BOM

Function remove_BOM

tools/tokenizer/token_common.c:85–100  ·  view source on GitHub ↗

Must be called right after a file is opened as stdin. Will attempt to remove any UTF-8 unicode signature (byte-order mark, BOM) at the beginning of the file. Unicode: U+FEFF UTF-8: EF BB BF First bytes Encoding Must remove? 00 00 FE FF UTF-32 big endian Yes FF FE 00 00 UTF-32 little endian Yes FE FF UTF-16 big endian Yes FF FE UTF-16 li

Source from the content-addressed store, hash-verified

83 otherwise UTF-8 No
84*/
85void remove_BOM(void)
86{
87 int c1 = getchar();
88 if (c1 == 0xEF) {
89 int c2 = getchar();
90 if (c2 == 0xBB) {
91 int c3 = getchar();
92 if (c3 == 0xBF) {
93 return;
94 }
95 if (c3 != EOF) buffer[buffered++] = c3;
96 }
97 if (c2 != EOF) buffer[buffered++] = c2;
98 }
99 if (c1 != EOF) buffer[buffered++] = c1;
100}
101
102/* Deal with DOS (\r \n) and classic Mac OS (\r) (physical) line endings.
103 In case of CR LF skip (but count) the CR and return LF.

Callers

nothing calls this directly

Calls

no outgoing calls

Tested by

no test coverage detected