Required prerequisites
Motivation
We work on vLLM-metal (https://github.com/vllm-project/vllm-metal), which uses MLX and has many handwritten Metal kernels. These kernels are difficult to write, review, and maintain.
I recently discovered TileLang's Metal DSL and would like to experiment with it as an alternative.
I tried a small local experiment on GDN in-place cache updates, and it indeed made the kernel code easier to express!!
However, the existing PyTorch Metal adapter does not directly fit our MLX runtime, so the experiment required custom MLX integration.
Would you be open to supporting an MLX adapter?
Solution
No response
Alternatives
No response
Additional context
No response
Required prerequisites
Motivation
We work on vLLM-metal (https://github.com/vllm-project/vllm-metal), which uses MLX and has many handwritten Metal kernels. These kernels are difficult to write, review, and maintain.
I recently discovered TileLang's Metal DSL and would like to experiment with it as an alternative.
I tried a small local experiment on GDN in-place cache updates, and it indeed made the kernel code easier to express!!
However, the existing PyTorch Metal adapter does not directly fit our MLX runtime, so the experiment required custom MLX integration.
Would you be open to supporting an MLX adapter?
Solution
No response
Alternatives
No response
Additional context
No response