| ▲ | tintor 5 hours ago | |||||||
Airgapped LLM inferrence server can't serve their output tokens, right? | ||||||||
| ▲ | angry_octet 4 hours ago | parent | next [-] | |||||||
They can expose just their inference port, possible via some supervisor. The inference consumer can also be air gapped. This kind of segmentation is increasingly common for high value services. | ||||||||
| ||||||||
| ▲ | bibimsz 4 hours ago | parent | prev [-] | |||||||
not at a high bitrate | ||||||||