Maps the words in the encoded file to their list of token ids. Reads all the subwords in encoded file. Concatenates them while they have @@ as their last two characters. The last token of a word is the subword without @@. Maps the full word to the list of corresponding vocab ids of the subwor
(encoded_filepath, vocab)