Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
Dylan16807
6 days ago
|
parent
|
context
|
favorite
| on:
“Next-token predictor” is the wrong mental model f...
You're using the fact the both parts of training affect the same weights to support your argument that they're making the system do something fundamentally
different
after RL?
help
danielmarkbruce
5 days ago
[–]
Assuming you are saying that RL is changing the model from doing one thing to another, yes. RL is changing the nature of the model.
reply
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: