| ▲ | xpct 8 hours ago | |
It's not intuitive to me for why preference for its own writing would emerge, and during what type of training or tuning. Perhaps something like: learning to identify what source files it has worked on by the code style alone, because tasks may give human code (public repos, etc) and ask to make changes. | ||
| ▲ | freeone3000 6 hours ago | parent | next [-] | |
It’s optimizing for good writing. Therefore, it believes its outputs are good. Therefore, it believes inputs that look like its outputs are good. | ||
| ▲ | pixl97 7 hours ago | parent | prev [-] | |
It would need to be researched, but I wonder if it ends up being something that happens at the token level? | ||