A byte-pair encoder for sub-word embeddings. Notes ----- Byte-pair encoding [1][2] is a compression algorithm that iteratively replaces the most frequently ocurring byte pairs in a set of documents with a new, single token. It has gained popularity a
(self, max_merges=3000, encoding="utf-8")
| 176 | |
| 177 | class BytePairEncoder(object): |
| 178 | def __init__(self, max_merges=3000, encoding="utf-8"): |
| 179 | """ |
| 180 | A byte-pair encoder for sub-word embeddings. |
| 181 | |
| 182 | Notes |
| 183 | ----- |
| 184 | Byte-pair encoding [1][2] is a compression algorithm that iteratively |
| 185 | replaces the most frequently ocurring byte pairs in a set of documents |
| 186 | with a new, single token. It has gained popularity as a preprocessing |
| 187 | step for many NLP tasks due to its simplicity and expressiveness: using |
| 188 | a base coebook of just 256 unique tokens (bytes), any string can be |
| 189 | encoded. |
| 190 | |
| 191 | References |
| 192 | ---------- |
| 193 | .. [1] Gage, P. (1994). A new algorithm for data compression. *C |
| 194 | Users Journal, 12(2)*, 23–38. |
| 195 | .. [2] Sennrich, R., Haddow, B., & Birch, A. (2015). Neural machine |
| 196 | translation of rare words with subword units, *Proceedings of the |
| 197 | 54th Annual Meeting of the Association for Computational |
| 198 | Linguistics,* 1715-1725. |
| 199 | |
| 200 | Parameters |
| 201 | ---------- |
| 202 | max_merges : int |
| 203 | The maximum number of byte pair merges to perform during the |
| 204 | :meth:`fit` operation. Default is 3000. |
| 205 | encoding : str |
| 206 | The encoding scheme for the documents used to train the encoder. |
| 207 | Default is `'utf-8'`. |
| 208 | """ |
| 209 | self.parameters = { |
| 210 | "max_merges": max_merges, |
| 211 | "encoding": encoding, |
| 212 | } |
| 213 | |
| 214 | # initialize the byte <-> token and token <-> byte dictionaries. bytes |
| 215 | # are represented in decimal as integers between 0 and 255. there is a |
| 216 | # 1:1 correspondence between token and byte representations up to 255. |
| 217 | self.byte2token = OrderedDict({i: i for i in range(256)}) |
| 218 | self.token2byte = OrderedDict({v: k for k, v in self.byte2token.items()}) |
| 219 | |
| 220 | def fit(self, corpus_fps, encoding="utf-8"): |
| 221 | """ |
nothing calls this directly
no outgoing calls
no test coverage detected