Downloads · 30 days
13
17% of all-time downloads
tstruk/under-fire-terrain-controlnet
under-fire-terrain-controlnet is a image-to-image model from tstruk. Use it when you need one image transformed into another. It is set up for diffusers. The card lists the license as creativeml-openrail-m.
A Stable Diffusion 1.5 ControlNet that turns annotated game-terrain renders from Under Fire (a free, open-development WW2 RTS) into photorealistic aerial terrain. Trained on tstruk/under-fire-terrain-pairs: 600 pairs…
Downloads · 30 days
13
17% of all-time downloads
All-time downloads
76
Public
Parameters
361M
5.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.3 GB · 100%
From the Hugging Face model README
A Stable Diffusion 1.5 ControlNet that turns annotated game-terrain renders from Under Fire (a free, open-development WW2 RTS) into photorealistic aerial terrain. Trained on tstruk/under-fire-terrain-pairs: 600 pairs of reference renders and gpt-image-2 photorealistic targets, cut into aligned 512x512 patches and split by map.
checkpoint-018000/: step 18,000 — softer, more photographic textures
than 30k with slightly weaker layout obedience; pairs well with a raised
controlnet_conditioning_scale (1.2-1.5).checkpoint-084000/: step 84,000 — the final checkpoint of the sweep.
Most layout-faithful but visibly overtrained on this dataset size (harsher,
oversharpened textures); published for comparison and study.The base model (VAE, UNet, text encoder) stays frozen SD 1.5; only the ControlNet was trained (bf16 activations, fp32 weights, batch 16 effective, lr 1e-5 warmup+cosine, noise-prediction MSE).
import torch
from diffusers import (ControlNetModel, StableDiffusionControlNetPipeline,
UniPCMultistepScheduler)
from PIL import Image
controlnet = ControlNetModel.from_pretrained(
"tstruk/under-fire-terrain-controlnet", torch_dtype=torch.bfloat16)
pipe = StableDiffusionControlNetPipeline.from_pretrained(
"stable-diffusion-v1-5/stable-diffusion-v1-5",
controlnet=controlnet, torch_dtype=torch.bfloat16, safety_checker=None)
pipe.scheduler = UniPCMultistepScheduler.from_config(pipe.scheduler.config)
pipe.to("cuda")
cond = Image.open("reference_512.png") # a 512x512 reference-render patch
img = pipe(
"ultra realistic aerial photograph of rural french countryside terrain, "
"drone photo, neutral daylight",
negative_prompt="blurry, low quality, oversharpened, noisy, grainy, "
"jpeg artifacts, painting, cartoon, illustration, oversaturated, "
"high contrast, hdr",
image=cond, num_inference_steps=25, guidance_scale=4.0,
controlnet_conditioning_scale=1.2,
).images[0]
controlnet_conditioning_scale: 1.0 default; raise to 1.2-1.5 to force
stricter terrain-type placement on the prettier checkpoints.The conditioning image must match the training distribution: the reference-render style documented on the dataset card (marker legend, ~21.6 px per world unit, camera tilt 35 degrees — the game's default tilt is 45 nowadays, so tilted conditioning should be captured at 35).
For whole game views rather than single patches: resize the frame so the ground scale matches training (~21.6 px per world unit, canonically 1718x915 per screen), crop into 512x512 tiles with ~110 px overlap, run the tiles as one batch, and stitch with linear blending across the overlaps. Two optional quality passes used in the game's own pipeline:
This model and its dataset exist thanks to Under Fire — a free, community-built WW2 real-time tactics game developed in the open, where this model bakes photorealistic terrain. Please credit underfire.io when building on this work.
Inherits CreativeML OpenRAIL-M from Stable Diffusion 1.5 (a ControlNet is initialized from, and contains weights derived from, the SD 1.5 UNet). The training dataset itself is CC0.