Downloads · 30 days
17
13% of all-time downloads
TensorAeroSpace/sac-b747
sac-b747 is a reinforcement learning model from TensorAeroSpace. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for tensoraerospace. The card lists the license as mit.
This model is a Soft Actor-Critic (SAC) agent trained to control the pitch channel of a Boeing 747 in the tensoraerospace.envs.b747.ImprovedB747Env environment. The agent tracks a reference pitch profile while minimiz…
Downloads · 30 days
17
13% of all-time downloads
All-time downloads
126
Public
Repo size
1.4 MB
Likes
0
Public
Click a slice to open those files.
.pth1.4 MB · 99%
From the Hugging Face model README
This model is a Soft Actor-Critic (SAC) agent trained to control the pitch channel of a Boeing 747 in the tensoraerospace.envs.b747.ImprovedB747Env environment. The agent tracks a reference pitch profile while minimizing control effort and promoting smoothness.
tensoraerospace.envs.b747.ImprovedB747Env[norm_pitch_error, norm_q, norm_theta, norm_prev_action]Use the pretrained policy for simulation of pitch tracking tasks in the provided environment. Suitable for research and demonstration of RL-based flight control.
ImprovedB747Env.pip install tensoraerospace
from tensoraerospace.agent.sac import SAC
agent = SAC.from_pretrained(
"./example/reinforcement_learning/best_episode_200k_episodes_0008_mae/Oct02_11-52-57_SAC/",
load_gradients=False, # set True to resume training with optimizer states
)
# Evaluate
obs, info = agent.env.reset()
done = False
while not done:
action = agent.select_action(obs, evaluate=True)
obs, reward, terminated, truncated, info = agent.env.step(action)
done = terminated or truncated
from tensoraerospace.agent.sac import SAC
agent = SAC.from_pretrained(
"./example/reinforcement_learning/best_episode_200k_episodes_0008_mae/Oct02_11-52-57_SAC/",
load_gradients=True,
)
agent.train(num_episodes=10)
agent.save("./runs", save_gradients=True)
The saved config.json contains the exact environment and policy parameters used for training. Key entries:
env.name: tensoraerospace.envs.b747.ImprovedB747Envenv.params:
initial_state: [0, 0, 0, 0]reference_signal: shape (1, 201) sinusoidal-like target for pitchnumber_time_steps: 201policy.params:
gamma: 0.99tau: 0.02alpha: auto via automatic entropy tuningbatch_size: 256updates_per_step: 2target_update_interval: 1lr: 3e-4policy_type: Gaussiandevice: cpuNote: With automatic_entropy_tuning=True, log_alpha and alpha_optim state are saved and can be restored.
The agent was validated in simulation on the same environment by tracking the provided reference pitch signal over 201 steps. Reward aligns with negative quadratic costs on tracking error, pitch rate, control magnitude, smoothness, and jerk.
Training performed on CPU for this checkpoint. For large-scale training, estimate CO2eq with the ML CO2 Impact calculator.
If you use this model, please cite the TensorAeroSpace repository.
@misc{tensoraerospace,
title = {TensorAeroSpace: Aerospace Simulation and RL Framework},
author = {TensorAeroSpace contributors},
year = {2023},
howpublished = {\url{https://github.com/tensoraerospace/tensoraerospace}},
}
TensorAeroSpace Team
For questions, please open an issue at the repository or email [email protected].