jdcsen Portfolio, projects, and other work by Joshua David Christensen
Projects with the tag vLLM

Paralinguistic Signals from a Speech LLM

  • NVIDIA Triton Python backend fronting a vLLM-served speech LLM, taken from prototype to production.
  • Derives confidence, sentiment and spoken-language ID from the model’s own token-level outputs. No additional models to train or serve.
  • Orchestrates per-signal LLM queries over gRPC alongside ONNX Runtime inference for confidence calibration and forced alignment.
  • p50 latency of approximately 5 to 10 ms per signal.
  • Technical lead for a three-engineer team. A later concurrency refactor took sustained throughput from roughly 8 to 80 requests per second.