Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
A simple fix for LLM tail latency (myhoai.com)
25 points by oskrim 3 hours ago | hide | past | favorite | 11 comments
 help



This sounds like a job for Fast Fallback instead: https://en.wikipedia.org/wiki/Happy_Eyeballs

Sending two identical parallel requests is the classic approach. But, logically speaking, it should also double the cost.

I would send a second request if the first request fails to return the first token within, say, 1 second. Then there's a chance the first request is stalling, which is an infrequent event.

I wonder if higher-availability tiers of LLM providers do a similar thing internally.


Token caching might help here, but if it returns the same result, faster, for the same price as priority, seems good

Nice turn around, does anyone has a benchmark regarding other types of requests (priority vs send twice) other than voice/call? Or the tests already test that?

I love this. Simple. Useful. To the point. If AI was used, I can't tell because it is clearly representing the author's beliefs.

agreed, reads like a breath of fresh air, no fluff

Why not send it thrice?

for a tier thats twice the cost i would expect >2x the speed. somewhere 5-10x

e.g. 1.40m would become 0.30s.

do people really pay for these priority plans?


This is half the story; you should show performance per dollar. I doubt your 2x approach would fare well against the priority if you consider the costs.

The article mentions that the priority tier costs 2x normal, so the costs of running normal twice should be fine.

If you want a controllable and predictable system, host it yourself. APIs will always have outages, delays and breaking changes every so often. That's the price you pay for not doing it properly and outsourcing your job.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: