MCPcopy Create free account
hub / github.com/CLUEbenchmark/CLUE / convert_single_example

Function convert_single_example

baselines/models/roberta/run_classifier.py:309–408  ·  view source on GitHub ↗

Converts a single `InputExample` into a single `InputFeatures`.

(ex_index, example, label_list, max_seq_length,
                           tokenizer)

Source from the content-addressed store, hash-verified

307
308
309def convert_single_example(ex_index, example, label_list, max_seq_length,
310 tokenizer):
311 """Converts a single `InputExample` into a single `InputFeatures`."""
312
313 if isinstance(example, PaddingInputExample):
314 return InputFeatures(
315 input_ids=[0] * max_seq_length,
316 input_mask=[0] * max_seq_length,
317 segment_ids=[0] * max_seq_length,
318 label_id=0,
319 is_real_example=False)
320
321 label_map = {}
322 for (i, label) in enumerate(label_list):
323 label_map[label] = i
324
325 tokens_a = tokenizer.tokenize(example.text_a)
326 tokens_b = None
327 if example.text_b:
328 tokens_b = tokenizer.tokenize(example.text_b)
329
330 if tokens_b:
331 # Modifies `tokens_a` and `tokens_b` in place so that the total
332 # length is less than the specified length.
333 # Account for [CLS], [SEP], [SEP] with "- 3"
334 _truncate_seq_pair(tokens_a, tokens_b, max_seq_length - 3)
335 else:
336 # Account for [CLS] and [SEP] with "- 2"
337 if len(tokens_a) > max_seq_length - 2:
338 tokens_a = tokens_a[0:(max_seq_length - 2)]
339
340 # The convention in BERT is:
341 # (a) For sequence pairs:
342 # tokens: [CLS] is this jack ##son ##ville ? [SEP] no it is not . [SEP]
343 # type_ids: 0 0 0 0 0 0 0 0 1 1 1 1 1 1
344 # (b) For single sequences:
345 # tokens: [CLS] the dog is hairy . [SEP]
346 # type_ids: 0 0 0 0 0 0 0
347 #
348 # Where "type_ids" are used to indicate whether this is the first
349 # sequence or the second sequence. The embedding vectors for `type=0` and
350 # `type=1` were learned during pre-training and are added to the wordpiece
351 # embedding vector (and position vector). This is not *strictly* necessary
352 # since the [SEP] token unambiguously separates the sequences, but it makes
353 # it easier for the model to learn the concept of sequences.
354 #
355 # For classification tasks, the first vector (corresponding to [CLS]) is
356 # used as the "sentence vector". Note that this only makes sense because
357 # the entire model is fine-tuned.
358 tokens = []
359 segment_ids = []
360 tokens.append("[CLS]")
361 segment_ids.append(0)
362 for token in tokens_a:
363 tokens.append(token)
364 segment_ids.append(0)
365 tokens.append("[SEP]")
366 segment_ids.append(0)

Calls 5

joinMethod · 0.80
InputFeaturesClass · 0.70
_truncate_seq_pairFunction · 0.70
tokenizeMethod · 0.45
convert_tokens_to_idsMethod · 0.45

Tested by

no test coverage detected