Guest post pitch for kernel.org: 3 ideas from a hands-on AI engineer

[email protected]
Newsgroups org.kernel.vger.stable
Message-ID <[email protected]>
A hard lesson from production: in a crypto accounting engine I worked on at Bitwave (Zealsight), we hit a scalability break where ingestion lag accumulated and downstream consumers fell behind. The fix wasn’t a single “optimization”; it was adding clearer queue/backlog visibility, tightening batching boundaries, and simplifying the data path so we could reason about throughput and failure behavior under load.

I’d like to pitch a guest post to kernel.org readers—grounded in measurement, hardware reality, and systems constraints—from my current work benchmarking local LLM inference across Apple Silicon, NVIDIA, and AMD (ROCm), plus prior distributed-systems work.

Possible titles (choose one):
1) Performance: “From Backlog to Baseline: A Practical Playbook for Measuring and Regressing Throughput in Production Pipelines”
2) GPU: “Local LLM Inference on Linux: What Benchmarks Reveal About NVIDIA vs ROCm vs Apple Silicon Constraints”
3) Distributed systems: “Designing Backpressure You Can Debug: Queue Signals, Batch Sizes, and Failure Containment”

Writing samples: kunalganglani.com

If you’d rather not hear from me, say so and I won’t email again.

Kunal
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.