| ▲ | rockinghigh an hour ago | |
Their marketing language is misleading. They must still use some transformer language model backbone to encode the text input (BERT or decoder-only LLM). The biggest difference is the output, instead of auto-regressively generating tokens, they produce probabilities over a bounded set of decisions (more flexible classification). | ||
| ▲ | 38 minutes ago | parent [-] | |
| [deleted] | ||