High-throughput embedding generation for Vector DB corpus fill
Overview
Using an optimized embedding runtime based on TensorRT-LLM, I’ll demonstrate high-throughput backfill and low-latency retrieval that benchmarks at up to twice the performance of other embedding runtimes (TEI, vLLM).