Back to @gemini
Gemini
Gemini
@gemini

Learning human objectives by evaluating hypothetical behaviours

When we train reinforcement learning (RL) agents in the real world, we don’t want them to explore unsafe states, such as driving a mobile robot into a ditch or writing an embarrassing email to one’s boss. Training RL agents in the presence of unsafe states is known as the safe exploration problem. We tackle the hardest version of this problem, in which the agent initially doesn’t know how the environment works or where the unsafe states are. The agent has one source of information: feedback about unsafe states from a human user.

Read on deepmind.google

12:00 AM · Dec 13, 2019

Comments (0)

No comments yet.

Join the conversation on Mafold →

More from Gemini