MCPcopy Create free account
hub / github.com/ddbourgin/numpy-ml / __init__

Method __init__

numpy_ml/preprocessing/nlp.py:178–218  ·  view source on GitHub ↗

A byte-pair encoder for sub-word embeddings. Notes ----- Byte-pair encoding [1][2] is a compression algorithm that iteratively replaces the most frequently ocurring byte pairs in a set of documents with a new, single token. It has gained popularity a

(self, max_merges=3000, encoding="utf-8")

Source from the content-addressed store, hash-verified

176
177class BytePairEncoder(object):
178 def __init__(self, max_merges=3000, encoding="utf-8"):
179 """
180 A byte-pair encoder for sub-word embeddings.
181
182 Notes
183 -----
184 Byte-pair encoding [1][2] is a compression algorithm that iteratively
185 replaces the most frequently ocurring byte pairs in a set of documents
186 with a new, single token. It has gained popularity as a preprocessing
187 step for many NLP tasks due to its simplicity and expressiveness: using
188 a base coebook of just 256 unique tokens (bytes), any string can be
189 encoded.
190
191 References
192 ----------
193 .. [1] Gage, P. (1994). A new algorithm for data compression. *C
194 Users Journal, 12(2)*, 23–38.
195 .. [2] Sennrich, R., Haddow, B., & Birch, A. (2015). Neural machine
196 translation of rare words with subword units, *Proceedings of the
197 54th Annual Meeting of the Association for Computational
198 Linguistics,* 1715-1725.
199
200 Parameters
201 ----------
202 max_merges : int
203 The maximum number of byte pair merges to perform during the
204 :meth:`fit` operation. Default is 3000.
205 encoding : str
206 The encoding scheme for the documents used to train the encoder.
207 Default is `'utf-8'`.
208 """
209 self.parameters = {
210 "max_merges": max_merges,
211 "encoding": encoding,
212 }
213
214 # initialize the byte <-> token and token <-> byte dictionaries. bytes
215 # are represented in decimal as integers between 0 and 255. there is a
216 # 1:1 correspondence between token and byte representations up to 255.
217 self.byte2token = OrderedDict({i: i for i in range(256)})
218 self.token2byte = OrderedDict({v: k for k, v in self.byte2token.items()})
219
220 def fit(self, corpus_fps, encoding="utf-8"):
221 """

Callers

nothing calls this directly

Calls

no outgoing calls

Tested by

no test coverage detected