Generate dense embedding vector for the input text. This method calls the OpenAI Embeddings API to convert input text into a dense vector representation. Results are cached to improve performance for repeated inputs. Args: input (TEXT): Input text string
(self, input: TEXT)
| 172 | |
| 173 | @lru_cache(maxsize=10) |
| 174 | def embed(self, input: TEXT) -> DenseVectorType: |
| 175 | """Generate dense embedding vector for the input text. |
| 176 | |
| 177 | This method calls the OpenAI Embeddings API to convert input text |
| 178 | into a dense vector representation. Results are cached to improve |
| 179 | performance for repeated inputs. |
| 180 | |
| 181 | Args: |
| 182 | input (TEXT): Input text string to embed. Must be non-empty after |
| 183 | stripping whitespace. Maximum length is 8191 tokens for most models. |
| 184 | |
| 185 | Returns: |
| 186 | DenseVectorType: A list of floats representing the embedding vector. |
| 187 | Length equals ``self.dimension``. Example: |
| 188 | ``[0.123, -0.456, 0.789, ...]`` |
| 189 | |
| 190 | Raises: |
| 191 | TypeError: If ``input`` is not a string. |
| 192 | ValueError: If input is empty/whitespace-only, or if the API returns |
| 193 | an error or malformed response. |
| 194 | RuntimeError: If network connectivity issues or OpenAI service |
| 195 | errors occur. |
| 196 | |
| 197 | Examples: |
| 198 | >>> emb = OpenAIDenseEmbedding() |
| 199 | >>> vector = emb.embed("Natural language processing") |
| 200 | >>> len(vector) |
| 201 | 1536 |
| 202 | >>> isinstance(vector[0], float) |
| 203 | True |
| 204 | |
| 205 | >>> # Error: empty input |
| 206 | >>> emb.embed(" ") |
| 207 | ValueError: Input text cannot be empty or whitespace only |
| 208 | |
| 209 | >>> # Error: non-string input |
| 210 | >>> emb.embed(123) |
| 211 | TypeError: Expected 'input' to be str, got int |
| 212 | |
| 213 | Note: |
| 214 | - This method is cached (maxsize=10). Identical inputs return cached results. |
| 215 | - The cache is based on exact string match (case-sensitive). |
| 216 | - Consider pre-processing text (lowercasing, normalization) for better caching. |
| 217 | """ |
| 218 | if not isinstance(input, TEXT): |
| 219 | raise TypeError(f"Expected 'input' to be str, got {type(input).__name__}") |
| 220 | |
| 221 | input = input.strip() |
| 222 | if not input: |
| 223 | raise ValueError("Input text cannot be empty or whitespace only") |
| 224 | |
| 225 | # Call API |
| 226 | embedding_vector = self._call_text_embedding_api( |
| 227 | input=input, |
| 228 | dimension=self._custom_dimension, |
| 229 | ) |
| 230 | |
| 231 | # Verify dimension |