clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption.
Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:
$ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose
[...]
File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name
raise ValueError(f"Can not map tensor {name!r}")
ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'