Difference between revisions of "Distributed Deep Reinforcement Learning (DDRL)"

From
Jump to: navigation, search
m (Text replacement - "http:" to "https:")
m
 
(One intermediate revision by the same user not shown)
Line 24: Line 24:
 
*** [[Lifelong Latent Actor-Critic (LILAC)]]
 
*** [[Lifelong Latent Actor-Critic (LILAC)]]
 
** [[Hierarchical Reinforcement Learning (HRL)]]
 
** [[Hierarchical Reinforcement Learning (HRL)]]
* [[Agents]]  ... [[Agents#Communication | communications]]
+
* [[Agents/Assistants]] ... [[Robotic Process Automation (RPA)|Robotic Process Automation]] ... [[Personal Companions]] ... [[Personal Productivity|Productivity]] ... [[Email]] ... [[Negotiation]] ... [[LangChain]]
 +
* [[Policy]]  ... [[Policy vs Plan]] ... [[Constitutional AI]] ... [[Trust Region Policy Optimization (TRPO)]] ... [[Policy Gradient (PG)]] ... [[Proximal Policy Optimization (PPO)]]
  
  
  
a new, highly scalable [[Agents|agent]] architecture for distributed training called Importance Weighted Actor-Learner Architecture that uses a new off-policy correction algorithm called V-trace.
+
a new, highly scalable [[Agents/Assistants|Agent]] architecture for distributed training called Importance Weighted Actor-Learner Architecture that uses a new off-policy correction algorithm called V-trace.
  
 
<youtube>-YMfJLFynmA</youtube>
 
<youtube>-YMfJLFynmA</youtube>

Latest revision as of 20:43, 17 September 2026