# UAV-DDPG: Deep Reinforcement Learning Code for UAV-Assisted Mobile Edge Computing

> UAV-DDPG is the reference Python implementation from a Wireless Networks paper that applies the Deep Deterministic Policy Gradient algorithm to minimize processing delay in a UAV-assisted mobile edge computing system. It is research code, not a general-purpose framework, but the environment and agent are documented in enough detail to reproduce the paper's results.

**fangvv/UAV-DDPG** — Code for paper "Computation Offloading Optimization for UAV-assisted Mobile Edge Computing: A Deep Deterministic Policy Gradient Approach"

- Repository: https://github.com/fangvv/UAV-DDPG
- Website: https://fangvv.github.io/UAV-DDPG/
- Stars: 722 · Forks: 95
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/fangvv-uav-ddpg

## The Research Problem: Joint Optimization of UAV Offloading

In a UAV-assisted mobile edge computing setup, a drone carrying a computing payload serves as an airborne edge server for ground user equipment. Each user can offload a fraction of its computing tasks to the UAV and execute the rest locally on its own processor. The system must decide, at each time step, which user to serve, how much of that user's task to offload, and where the UAV should fly next. The objective is to minimize the maximum processing delay across all users.

The difficulty is that these decisions are continuous, coupled, and taken under uncertainty. The flight angle, the flight distance, and the offloading ratio are all real-valued. The UAV's position affects the wireless channel quality, which in turn affects transmission delay, which feeds back into the offloading decision. Deep Q-Network algorithms require discretizing the action space, which introduces approximation error in continuous control problems. DDPG operates directly on continuous actions by maintaining a deterministic policy, which is why it is the algorithm of choice in this paper.

## Environment Model in UAVEnv

The simulation environment is implemented in `DDPG/UAV_env.py`. The modeled area is a 100-meter cube. The channel bandwidth is set to 1 MHz and the UAV flight speed is 50 m/s. With these parameters the system runs for 40 time slots per episode.

The state vector has dimension `4 + M x 4`, where M is the number of user equipments (set to 4 by default). It carries the UAV battery level, UAV position, remaining task size, per-user locations, per-user task sizes, and a line-of-sight flag for each user. The action vector is four-dimensional: a continuous value mapped to a target user index, a flight angle in `[0, 2pi]`, a flight distance bounded by speed and time, and a task offloading ratio in `[0, 1]`.

The reward is the negative of the maximum processing delay across all users at each step. The `com_delay` method computes this delay as the maximum of transmission delay plus edge computation time versus local computation time. The channel model follows free-space path loss with separate noise levels for line-of-sight and non-line-of-sight conditions.

## Installing and Running the Code

The repository has no package configuration. It runs with TensorFlow 1.X, NumPy, and Matplotlib. It does not install via pip; you clone the repository and run the scripts directly. The required software list in the README is:

- TensorFlow 1.X
- NumPy
- Matplotlib

The project structure places the main DDPG training code in the `DDPG/` directory and the DQN baseline in `DQN/`. Ablation variants that remove exploration noise or state normalization are sub-directories inside `DDPG/`:

```
UAV-DDPG/
├── DDPG/
│   ├── UAV_env.py
│   ├── ddpg_algo.py
│   ├── state_normalization.py
│   ├── DDPG_without_behavior_noise/
│   └── DDPG_without_state_normalization/
├── DQN/
├── Edge_only/
├── Local_only/
└── README.md
```

To run the DDPG training, enter the `DDPG/` directory and execute the training script with Python. No command-line arguments are documented in the README; hyperparameters are set in `ddpg_algo.py` directly.

## DDPG Agent Architecture and Hyperparameters

The agent is implemented in `DDPG/ddpg_algo.py`. Both the Actor and Critic networks use three hidden layers. The Actor maps the state through layers of width 400, 300, and 10, with ReLU6 activations on the first two and a final tanh scaled by `action_bound`. The Critic embeds the state and the action separately into 400 units, sums them with a bias, then passes the result through layers of width 300 and 10 before outputting a scalar Q-value.

Experience replay uses a buffer of 10,000 transitions and a batch size of 64. Soft target updates use tau = 0.01, meaning target network weights move 1% toward the online network each training step. The discount factor gamma is 0.001, which is unusual: it weights the immediate reward heavily relative to future rewards, a deliberate choice for a short-horizon delay minimization problem. The learning rates are 0.001 for the Actor and 0.002 for the Critic.

During training, Gaussian noise is added to the actor's output to encourage exploration. The `state_normalization.py` module divides each state component by its maximum possible value, scaling everything to `[0, 1]`. For example, UAV battery is normalized by 500,000 joules, UAV location by 100 meters, and remaining task size by approximately 100 megabytes. The ablation variants in the sub-directories let a reader isolate the contribution of each component, comparing training with and without exploration noise or state normalization.

## DQN Baseline and Comparison

The `DQN/` directory contains a baseline agent that discretizes the continuous action space into `M x 11^3` discrete actions. With M=4 users and 11 steps on each of the three continuous dimensions, that yields 5,324 discrete actions. The paper reports that DDPG achieves a significantly lower maximum processing delay than this DQN baseline, and the repository structure makes it straightforward to run both agents against the same environment and compare their delay curves.

The `Edge_only/` and `Local_only/` directories serve as additional baselines: one routes all computation to the UAV, the other executes everything locally on the user equipment. These bound the achievable delay from above and provide reference points for evaluating the DDPG policy's improvement relative to naive strategies.

DDPG is the better match for this problem because the action space is inherently continuous. Discretizing flight angle and offloading ratio into 11 levels each introduces quantization error that is absent in the DDPG formulation. For other edge computing problems where the action space can be naturally discretized (for example, fixed-rate scheduling on a small number of channels), DQN can be competitive without the policy gradient overhead. The existence of both implementations in a single repository makes UAV-DDPG useful as a teaching example of why algorithm choice matters for continuous control.

## Limitations and Scope

The code implements one specific scenario from one paper. The environment is a fixed simulation: the area dimensions, channel model parameters, number of users, and timing structure are all hardcoded. There is no configuration file or command-line interface for changing them; a researcher who wants to modify the scenario must edit the source files.

TensorFlow 1.X is required. TensorFlow 2.X changed the default execution model and is not compatible with the session-based code in this repository. Python environments with TensorFlow 2.X installed will fail to run the code without migration.

The repository has no releases and no versioning. The last push was on 2026-07-22. There is no inference or deployment pipeline: the code trains a policy and logs convergence curves; it does not produce a deployable controller for a real UAV.

Researchers who want a maintained MEC simulation framework with configurable topologies and standard benchmarks will need to look elsewhere. UAV-DDPG is a single-paper artifact, valuable for the specific reproducibility purpose it was created for. The ablation variants in the sub-directories are its most useful feature beyond the main training run, because they let a reader test how much each component contributes to convergence without re-implementing the experiment from scratch.

## Conclusion

UAV-DDPG is useful for researchers who want to replicate or extend the paper's UAV-assisted MEC offloading results. The code structure is clean and the environment model is documented at the level needed to modify it. It is not designed for production deployment or as a general reinforcement learning library: TensorFlow 1.X is a hard requirement, the environment is a fixed simulation with no support for real hardware, and the codebase has no releases. Before using it, confirm that your Python environment has TensorFlow 1.X installed; later versions are not compatible with the session-based code.

## FAQ

### What is the Deep Deterministic Policy Gradient algorithm used in UAV-DDPG?

DDPG is a reinforcement learning algorithm for continuous action spaces. It maintains an Actor network that selects actions deterministically and a Critic network that estimates their Q-value, training both using experience replay and soft target network updates.

### Is DDPG a form of Q-learning?

DDPG uses a Critic network that approximates Q-values, which connects it to the Q-learning family. The Critic is trained with a temporal-difference target similar to DQN, but the Actor is trained by policy gradient rather than by argmax over the action space.

### What TensorFlow version does UAV-DDPG require?

The README specifies TensorFlow 1.X. The session-based code is not compatible with TensorFlow 2.X without migration.

## Sources

- [fangvv/UAV-DDPG on GitHub](https://github.com/fangvv/UAV-DDPG)
- [Issues](https://github.com/fangvv/UAV-DDPG/issues)
- [Project website](https://fangvv.github.io/UAV-DDPG/)
- [README](https://github.com/fangvv/UAV-DDPG/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fangvv-uav-ddpg
