Compare commits

..
9 Commits
Author SHA1 Message Date
twothatITandClaude Sonnet 5 0e2b292408 feat(customers): sortable table columns + default ascending ID order
- Customer list now defaults to ascending ID order instead of newest-first,
  so the table starts at customer #1 instead of the highest ID
- Add sort_by/sort_order query params to GET /customers (whitelisted column
  map to prevent SQL injection via arbitrary column names)
- Make ID/Name/Subdomain/Status/Devices/Created column headers clickable,
  toggling asc/desc with a visual arrow indicator

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-07-23 15:16:31 +02:00
twothatITandClaude Sonnet 5 a5988af6a3 fix(update): stop blocking event loop during rebuild + fix infinite spinner
The update endpoint ran the entire git pull + docker build (up to 10 min)
synchronously inside the request handler, blocking the whole server for
everyone while it ran. Separately, the frontend spinner was only hidden on
error, never on success, so it spun forever even when the update worked.

- Run the update in a background thread; the request returns immediately
- Add GET /settings/update/status for progress polling (backup/pull/build/restart)
- Frontend polls status, then waits for the app to come back after the
  container restart, and shows a clear done/timeout message instead of an
  endless spinner

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-07-23 15:09:42 +02:00
twothatITandClaude Sonnet 5 c5189d88fe perf(npm): cache NPM JWT instead of re-authenticating on every API call
Every NPM helper (proxy host create/update/delete, streams, certs) did a
fresh POST /api/tokens login before its actual request, adding an avoidable
round-trip to every proxy/stream operation.

- Cache the JWT per (api_url, email), sized from its 'exp' claim
- Transparently re-authenticate and retry once on a 401 (e.g. after an NPM
  restart invalidates a cached token), so a stale cache entry can't cause a
  hard failure

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-07-23 14:51:23 +02:00
twothatITandClaude Sonnet 5 ac843da4ca perf(monitoring): stop blocking event loop with synchronous Docker calls
Customer search and detail loads were intermittently slow because every
customer-table render (including each search keystroke) triggered
/monitoring/customers/local-update-status, which looped synchronously over
all customers doing blocking `docker inspect` subprocess calls on the event
loop — stalling all other in-flight requests, including search itself.

- Offload per-service image/container inspection to the thread pool and run
  checks concurrently instead of sequentially (image_service, docker_service)
- Reuse a single Docker SDK client instead of reconnecting per customer
- Cache local-update-status results for 20s since the underlying data only
  changes after an image pull, not on every keystroke
- Parallelize /monitoring/customers container status lookups

Co-Authored-By: Claude Sonnet 5 <[email protected]>
2026-07-23 14:44:56 +02:00
twothatITandClaude Sonnet 4.6 f6b7eb2dae fix(npm): add gRPC read/send timeouts to proxy host location blocks
Adds grpc_read_timeout 3600s and grpc_send_timeout 3600s to both
ManagementService and SignalExchange location blocks to prevent
long-lived gRPC connections from being dropped by Nginx.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
2026-05-06 12:01:14 +02:00
twothatITandClaude Sonnet 4.6 8ede0f0a3c fix(deploy): fix redeploy button broken by JSON.stringify double quotes
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
2026-03-10 22:13:23 +01:00
twothatITandClaude Sonnet 4.6 8040973227 fix(deploy): fix redeploy button broken by JSON.stringify double quotes
JSON.stringify('Name') produces "Name" with double quotes which breaks
the onclick attribute. Use data-customer-name attribute instead and
read it via this.dataset.customerName to avoid quoting issues.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
2026-03-10 22:13:06 +01:00
twothatITandClaude Sonnet 4.6 3cdc82f919 fix(deploy): show customer name in redeploy modal instead of ID
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
2026-03-10 22:08:35 +01:00
twothatITandClaude Sonnet 4.6 40595fc381 fix(deploy): show customer name in redeploy modal instead of ID
The modal was showing '#2' instead of the customer name when opened
from the customer detail view, because the dashboard table row was
not visible. Now the name is passed directly from the button's onclick
context where data.name is already available.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
2026-03-10 22:08:17 +01:00
13 changed files with 432 additions and 55 deletions
+1 -1
View File
@@ -33,7 +33,7 @@ logger = logging.getLogger(__name__)
app = FastAPI(
title="NetBird MSP Appliance",
description="Multi-tenant NetBird management platform for MSPs",
version="1.0.0",
version="1.2.0",
docs_url="/api/docs",
redoc_url="/api/redoc",
openapi_url="/api/openapi.json",
+19 -2
View File
@@ -78,22 +78,36 @@ async def create_customer(
return response
SORTABLE_CUSTOMER_COLUMNS = {
"id": Customer.id,
"name": Customer.name,
"subdomain": Customer.subdomain,
"status": Customer.status,
"max_devices": Customer.max_devices,
"created_at": Customer.created_at,
}
@router.get("")
async def list_customers(
page: int = Query(default=1, ge=1),
per_page: int = Query(default=25, ge=1, le=100),
search: Optional[str] = Query(default=None),
status_filter: Optional[str] = Query(default=None, alias="status"),
sort_by: str = Query(default="id"),
sort_order: str = Query(default="asc", pattern="^(asc|desc)$"),
current_user: User = Depends(get_current_user),
db: Session = Depends(get_db),
):
"""List customers with pagination, search, and status filter.
"""List customers with pagination, search, status filter, and sorting.
Args:
page: Page number (1-indexed).
per_page: Items per page.
search: Search in name, subdomain, email.
status_filter: Filter by status.
sort_by: Column to sort by — one of SORTABLE_CUSTOMER_COLUMNS.
sort_order: "asc" or "desc".
Returns:
Paginated customer list with metadata.
@@ -113,8 +127,11 @@ async def list_customers(
query = query.filter(Customer.status == status_filter)
total = query.count()
sort_column = SORTABLE_CUSTOMER_COLUMNS.get(sort_by, Customer.id)
sort_expr = sort_column.desc() if sort_order == "desc" else sort_column.asc()
customers = (
query.order_by(Customer.created_at.desc())
query.order_by(sort_expr, Customer.id.asc())
.offset((page - 1) * per_page)
.limit(per_page)
.all()
+32 -9
View File
@@ -1,7 +1,9 @@
"""Monitoring API — system overview, customer statuses, host resources."""
import asyncio
import logging
import platform
import time
from typing import Any
import psutil
@@ -16,6 +18,13 @@ from app.services import docker_service, image_service
logger = logging.getLogger(__name__)
router = APIRouter()
# Short-lived cache for the local update-status badges. This endpoint is
# triggered on every customer-table render (i.e. every search keystroke), but
# the underlying data (which images are outdated) only changes after an image
# pull + container recreate, so a few seconds of staleness is harmless.
_update_status_cache: dict[str, Any] = {"data": None, "expires": 0.0}
_UPDATE_STATUS_TTL_SECONDS = 20
@router.get("/status")
async def system_status(
@@ -58,8 +67,7 @@ async def all_customers_status(
.all()
)
results: list[dict[str, Any]] = []
for c in customers:
async def _build_entry(c: Customer) -> dict[str, Any]:
entry: dict[str, Any] = {
"id": c.id,
"name": c.name,
@@ -67,7 +75,7 @@ async def all_customers_status(
"status": c.status,
}
if c.deployment:
containers = docker_service.get_container_status(c.deployment.container_prefix)
containers = await docker_service.get_container_status_async(c.deployment.container_prefix)
entry["deployment_status"] = c.deployment.deployment_status
entry["containers"] = containers
entry["relay_udp_port"] = c.deployment.relay_udp_port
@@ -76,9 +84,11 @@ async def all_customers_status(
else:
entry["deployment_status"] = None
entry["containers"] = []
results.append(entry)
return entry
return results
# Fetch container status for all customers concurrently instead of one
# blocking Docker SDK call at a time.
return await asyncio.gather(*[_build_entry(c) for c in customers])
@router.get("/resources")
@@ -205,15 +215,28 @@ async def customers_local_update_status(
Compares running container image IDs against locally stored images.
No network call — safe to call on every dashboard load.
Results are cached for a few seconds since this is triggered on every
customer-table render (including every search keystroke) but the
underlying data rarely changes.
"""
now = time.monotonic()
if _update_status_cache["data"] is not None and now < _update_status_cache["expires"]:
return _update_status_cache["data"]
config = db.query(SystemConfig).filter(SystemConfig.id == 1).first()
if not config:
return []
deployments = db.query(Deployment).all()
results = []
for dep in deployments:
cs = image_service.get_customer_container_image_status(dep.container_prefix, config)
results.append({"customer_id": dep.customer_id, "needs_update": cs["needs_update"]})
async def _check(dep: Deployment) -> dict[str, Any]:
cs = await image_service.get_customer_container_image_status_async(dep.container_prefix, config)
return {"customer_id": dep.customer_id, "needs_update": cs["needs_update"]}
results = await asyncio.gather(*[_check(dep) for dep in deployments])
results = list(results)
_update_status_cache["data"] = results
_update_status_cache["expires"] = now + _UPDATE_STATUS_TTL_SECONDS
return results
+39 -10
View File
@@ -4,6 +4,7 @@ There is no .env file. Every setting lives in the ``system_config`` table
(singleton row with id=1) and is editable via the Web UI settings page.
"""
import asyncio
import logging
import os
import shutil
@@ -358,13 +359,15 @@ async def trigger_update(
current_user: User = Depends(get_current_user),
db: Session = Depends(get_db),
):
"""Backup the database, git pull the latest code, and rebuild the container.
"""Kick off backup + git pull + container rebuild in the background.
Returns immediately — the actual work (which can take several minutes,
especially the ``--no-cache`` image build) runs in a background thread so
it doesn't block this request or any other user's requests while it
runs. Progress can be polled via GET /settings/update/status until the
container restarts with the new version.
The rebuild is fire-and-forget — the app will restart in ~60 seconds.
Only admin users may trigger an update.
Returns:
Dict with ok, message, and backup path.
"""
if getattr(current_user, "role", "admin") != "admin":
raise HTTPException(
@@ -383,11 +386,37 @@ async def trigger_update(
detail="git_repo_url is not configured in settings.",
)
result = update_service.trigger_update(config, DATABASE_PATH)
if not result.get("ok"):
current_status = update_service.get_update_status()
if current_status.get("state") == "running":
raise HTTPException(
status_code=status.HTTP_500_INTERNAL_SERVER_ERROR,
detail=result.get("message", "Update failed."),
status_code=status.HTTP_409_CONFLICT,
detail="An update is already in progress.",
)
# Snapshot the only fields trigger_update() needs — avoids passing a
# SQLAlchemy instance into a background thread after this request's
# session may already be closed.
class _ConfigSnapshot:
git_repo_url = config.git_repo_url
git_branch = config.git_branch
git_token = config.git_token
asyncio.create_task(asyncio.to_thread(update_service.trigger_update, _ConfigSnapshot(), DATABASE_PATH))
logger.info("Update triggered by %s.", current_user.username)
return result
return {
"ok": True,
"message": "Update gestartet. Dies kann mehrere Minuten dauern — Fortschritt via Status sichtbar.",
}
@router.get("/update/status")
async def update_status(
current_user: User = Depends(get_current_user),
):
"""Return progress of the currently running (or last) update.
Note: once the container restarts mid-update, this endpoint stops
responding for a few seconds — that itself is a signal the swap is
happening. The frontend falls back to polling for the app coming back up.
"""
return update_service.get_update_status()
+23 -2
View File
@@ -27,13 +27,24 @@ async def _run_cmd(cmd: list[str], timeout: int = 120) -> subprocess.CompletedPr
)
_client: Optional[docker.DockerClient] = None
def _get_client() -> docker.DockerClient:
"""Return a Docker client connected via the Unix socket.
"""Return a shared Docker client connected via the Unix socket.
The client is created once and reused — creating a new client per call
(as `docker.from_env()` does) re-negotiates the API version and opens a
fresh connection every time, which is wasteful when called once per
customer in a loop.
Returns:
docker.DockerClient instance.
"""
return docker.from_env()
global _client
if _client is None:
_client = docker.from_env()
return _client
async def compose_up(
@@ -212,6 +223,16 @@ def get_container_status(container_prefix: str) -> list[dict[str, Any]]:
return results
async def get_container_status_async(container_prefix: str) -> list[dict[str, Any]]:
"""Thread-offloaded wrapper around get_container_status().
Use this when checking status for multiple customers so the Docker SDK
calls run in the thread pool instead of blocking the event loop.
"""
loop = asyncio.get_event_loop()
return await loop.run_in_executor(None, get_container_status, container_prefix)
def get_container_logs(container_name: str, tail: int = 200) -> str:
"""Retrieve recent logs from a container.
+38
View File
@@ -211,6 +211,44 @@ async def pull_all_images(config) -> dict[str, Any]:
}
async def get_customer_container_image_status_async(container_prefix: str, config) -> dict[str, Any]:
"""Async, thread-offloaded version of get_customer_container_image_status().
Runs the per-service `docker inspect` subprocess calls concurrently in the
thread pool instead of sequentially blocking the event loop — use this
whenever checking status for multiple customers (e.g. dashboard/search
badge refresh, monitoring overview).
Returns:
services: dict mapping service name to status info
needs_update: True if any service has a different image ID than locally stored
"""
service_images = {
"management": config.netbird_management_image,
"signal": config.netbird_signal_image,
"relay": config.netbird_relay_image,
"dashboard": config.netbird_dashboard_image,
}
loop = asyncio.get_event_loop()
async def _check(svc: str, image: str) -> tuple[str, dict[str, Any]]:
container_name = f"{container_prefix}-{svc}"
container_id, local_id = await asyncio.gather(
loop.run_in_executor(None, get_container_image_id, container_name),
loop.run_in_executor(None, get_local_image_id, image),
)
if container_id and local_id:
up_to_date = container_id == local_id
else:
up_to_date = None # container not running or image not pulled
return svc, {"container": container_name, "image": image, "up_to_date": up_to_date}
pairs = await asyncio.gather(*[_check(svc, image) for svc, image in service_images.items()])
services = dict(pairs)
needs_update = any(s["up_to_date"] is False for s in services.values())
return {"services": services, "needs_update": needs_update}
def get_customer_container_image_status(container_prefix: str, config) -> dict[str, Any]:
"""Check which service containers are running outdated local images.
+94 -16
View File
@@ -12,9 +12,12 @@ Let's Encrypt SSL certificates.
Also manages NPM streams for STUN/TURN relay UDP ports.
"""
import base64
import json
import logging
import os
import socket
import time
from typing import Any
import httpx
@@ -24,6 +27,14 @@ logger = logging.getLogger(__name__)
# Timeout for NPM API calls (seconds)
NPM_TIMEOUT = 30
# Cached JWTs, keyed by (api_url, email). NPM issues a token that stays valid
# for a while (per its 'exp' claim), so re-logging in on every single API
# call — as this module used to do — adds a full extra round-trip per action
# for no reason.
_token_cache: dict[tuple[str, str], dict[str, Any]] = {}
_TOKEN_SAFETY_MARGIN = 60 # refresh this many seconds before actual expiry
_DEFAULT_TOKEN_TTL = 3600 # fallback if the 'exp' claim can't be parsed
def _get_forward_host() -> str:
"""Get the host machine's real IP address for NPM forwarding.
@@ -90,6 +101,61 @@ async def _npm_login(client: httpx.AsyncClient, api_url: str, email: str, passwo
)
def _decode_jwt_exp(token: str) -> float | None:
"""Best-effort decode of a JWT's 'exp' claim, without verifying the signature.
We only use this to size our own cache TTL — NPM itself still enforces
the real expiry server-side, so an inaccurate read here is harmless.
"""
try:
payload_b64 = token.split(".")[1]
padding = "=" * (-len(payload_b64) % 4)
payload = json.loads(base64.urlsafe_b64decode(payload_b64 + padding))
return payload.get("exp")
except Exception:
return None
async def _get_token(
client: httpx.AsyncClient, api_url: str, email: str, password: str, force_refresh: bool = False
) -> str:
"""Return a cached NPM JWT if still valid, otherwise log in and cache it."""
cache_key = (api_url, email)
if not force_refresh:
cached = _token_cache.get(cache_key)
if cached and time.time() < cached["expires_at"]:
return cached["token"]
token = await _npm_login(client, api_url, email, password)
exp = _decode_jwt_exp(token)
expires_at = (exp - _TOKEN_SAFETY_MARGIN) if exp else (time.time() + _DEFAULT_TOKEN_TTL)
_token_cache[cache_key] = {"token": token, "expires_at": expires_at}
return token
async def _request_with_reauth(
client: httpx.AsyncClient,
method: str,
api_url: str,
email: str,
password: str,
path: str,
headers: dict,
**kwargs: Any,
) -> tuple[httpx.Response, dict]:
"""Perform a request; if the cached token was rejected, refresh and retry once.
Returns the response and the (possibly updated) headers dict, so callers
can reuse the fresh token for any further requests in the same session.
"""
resp = await client.request(method, f"{api_url}{path}", headers=headers, **kwargs)
if resp.status_code == 401:
token = await _get_token(client, api_url, email, password, force_refresh=True)
headers = {**headers, "Authorization": f"Bearer {token}"}
resp = await client.request(method, f"{api_url}{path}", headers=headers, **kwargs)
return resp, headers
async def test_npm_connection(api_url: str, email: str, password: str) -> dict[str, Any]:
"""Test connectivity to NPM by logging in and listing proxy hosts.
@@ -103,9 +169,11 @@ async def test_npm_connection(api_url: str, email: str, password: str) -> dict[s
"""
try:
async with httpx.AsyncClient(timeout=NPM_TIMEOUT) as client:
token = await _npm_login(client, api_url, email, password)
token = await _get_token(client, api_url, email, password)
headers = {"Authorization": f"Bearer {token}"}
resp = await client.get(f"{api_url}/nginx/proxy-hosts", headers=headers)
resp, headers = await _request_with_reauth(
client, "GET", api_url, email, password, "/nginx/proxy-hosts", headers
)
if resp.status_code == 200:
count = len(resp.json())
return {"ok": True, "message": f"Connected. Login OK. {count} proxy hosts found."}
@@ -136,9 +204,11 @@ async def list_certificates(api_url: str, email: str, password: str) -> dict[str
"""
try:
async with httpx.AsyncClient(timeout=NPM_TIMEOUT) as client:
token = await _npm_login(client, api_url, email, password)
token = await _get_token(client, api_url, email, password)
headers = {"Authorization": f"Bearer {token}"}
resp = await client.get(f"{api_url}/nginx/certificates", headers=headers)
resp, headers = await _request_with_reauth(
client, "GET", api_url, email, password, "/nginx/certificates", headers
)
if resp.status_code == 200:
result = []
for cert in resp.json():
@@ -263,10 +333,14 @@ async def create_proxy_host(
"location ^~ /management.ManagementService/ {\n"
f" grpc_pass grpc://{forward_host}:{forward_port};\n"
" grpc_set_header Host $host;\n"
" grpc_read_timeout 3600s;\n"
" grpc_send_timeout 3600s;\n"
"}\n"
"location ^~ /signalexchange.SignalExchange/ {\n"
f" grpc_pass grpc://{forward_host}:{forward_port};\n"
" grpc_set_header Host $host;\n"
" grpc_read_timeout 3600s;\n"
" grpc_send_timeout 3600s;\n"
"}\n"
),
"meta": {
@@ -278,14 +352,15 @@ async def create_proxy_host(
try:
async with httpx.AsyncClient(timeout=180) as client: # Long timeout for LE cert
token = await _npm_login(client, api_url, npm_email, npm_password)
token = await _get_token(client, api_url, npm_email, npm_password)
headers = {
"Authorization": f"Bearer {token}",
"Content-Type": "application/json",
}
resp = await client.post(
f"{api_url}/nginx/proxy-hosts", json=payload, headers=headers
resp, headers = await _request_with_reauth(
client, "POST", api_url, npm_email, npm_password,
"/nginx/proxy-hosts", headers, json=payload,
)
if resp.status_code in (200, 201):
data = resp.json()
@@ -538,14 +613,15 @@ async def create_stream(
try:
async with httpx.AsyncClient(timeout=NPM_TIMEOUT) as client:
token = await _npm_login(client, api_url, npm_email, npm_password)
token = await _get_token(client, api_url, npm_email, npm_password)
headers = {
"Authorization": f"Bearer {token}",
"Content-Type": "application/json",
}
resp = await client.post(
f"{api_url}/nginx/streams", json=payload, headers=headers
resp, headers = await _request_with_reauth(
client, "POST", api_url, npm_email, npm_password,
"/nginx/streams", headers, json=payload,
)
if resp.status_code in (200, 201):
data = resp.json()
@@ -583,10 +659,11 @@ async def delete_stream(
"""
try:
async with httpx.AsyncClient(timeout=NPM_TIMEOUT) as client:
token = await _npm_login(client, api_url, npm_email, npm_password)
token = await _get_token(client, api_url, npm_email, npm_password)
headers = {"Authorization": f"Bearer {token}"}
resp = await client.delete(
f"{api_url}/nginx/streams/{stream_id}", headers=headers
resp, headers = await _request_with_reauth(
client, "DELETE", api_url, npm_email, npm_password,
f"/nginx/streams/{stream_id}", headers,
)
if resp.status_code in (200, 204):
logger.info("Deleted NPM stream %d", stream_id)
@@ -619,10 +696,11 @@ async def delete_proxy_host(
"""
try:
async with httpx.AsyncClient(timeout=NPM_TIMEOUT) as client:
token = await _npm_login(client, api_url, npm_email, npm_password)
token = await _get_token(client, api_url, npm_email, npm_password)
headers = {"Authorization": f"Bearer {token}"}
resp = await client.delete(
f"{api_url}/nginx/proxy-hosts/{proxy_id}", headers=headers
resp, headers = await _request_with_reauth(
client, "DELETE", api_url, npm_email, npm_password,
f"/nginx/proxy-hosts/{proxy_id}", headers,
)
if resp.status_code in (200, 204):
logger.info("Deleted NPM proxy host %d", proxy_id)
+42
View File
@@ -20,6 +20,32 @@ SERVICE_NAME = "netbird-msp-appliance"
logger = logging.getLogger(__name__)
# In-memory progress tracker for the currently running (or last) update.
# The container gets replaced mid-update, so this deliberately does NOT need
# to survive a restart — the frontend detects completion by polling until the
# app comes back up and reports a new version, not by reading a final status
# here. It exists so the UI can show *something* other than a frozen spinner
# while the backup/pull/build steps are in progress.
_update_status: dict[str, Any] = {
"state": "idle", # idle | running | failed
"step": "",
"message": "",
"started_at": None,
}
def get_update_status() -> dict[str, Any]:
"""Return a snapshot of the current update progress."""
return dict(_update_status)
def _set_status(state: str, step: str, message: str = "") -> None:
_update_status["state"] = state
_update_status["step"] = step
_update_status["message"] = message
if step == "backup":
_update_status["started_at"] = datetime.utcnow().isoformat()
def _get_compose_project_name() -> str:
"""Detect the compose project name from the running container's labels.
@@ -233,10 +259,12 @@ def trigger_update(config: Any, db_path: str) -> dict:
Dict with ok (bool), message, backup path, and pulled_branch.
"""
# 1. Backup database before any changes
_set_status("running", "backup", "Datenbank wird gesichert …")
try:
backup_path = backup_database(db_path)
except Exception as exc:
logger.error("Database backup failed: %s", exc)
_set_status("failed", "backup", f"Database backup failed: {exc}")
return {"ok": False, "message": f"Database backup failed: {exc}", "backup": None}
# 2. Build git pull command (embed token in URL if provided)
@@ -252,6 +280,7 @@ def trigger_update(config: Any, db_path: str) -> dict:
pull_cmd = ["git", "-C", SOURCE_DIR, "pull", "origin", branch]
# 3. Git pull (synchronous — must complete before rebuild)
_set_status("running", "pull", f"Code wird von Branch '{branch}' geholt …")
# Ensure .git directory is owned by the process user (root inside container).
# The .git dir may be owned by the host user after manual operations.
try:
@@ -270,13 +299,16 @@ def trigger_update(config: Any, db_path: str) -> dict:
timeout=120,
)
except subprocess.TimeoutExpired:
_set_status("failed", "pull", "git pull timed out after 120s.")
return {"ok": False, "message": "git pull timed out after 120s.", "backup": backup_path}
except Exception as exc:
_set_status("failed", "pull", f"git pull error: {exc}")
return {"ok": False, "message": f"git pull error: {exc}", "backup": backup_path}
if result.returncode != 0:
stderr = result.stderr.strip()[:500]
logger.error("git pull failed (exit %d): %s", result.returncode, stderr)
_set_status("failed", "pull", f"git pull failed: {stderr}")
return {
"ok": False,
"message": f"git pull failed: {stderr}",
@@ -348,6 +380,7 @@ def trigger_update(config: Any, db_path: str) -> dict:
SERVICE_NAME,
]
logger.info("Phase A: building new image …")
_set_status("running", "build", "Docker-Image wird gebaut (kann mehrere Minuten dauern) …")
try:
build_result = subprocess.run(
build_cmd,
@@ -360,15 +393,18 @@ def trigger_update(config: Any, db_path: str) -> dict:
f.write(build_result.stderr)
if build_result.returncode != 0:
logger.error("Image build failed: %s", build_result.stderr[:500])
_set_status("failed", "build", f"Image build failed: {build_result.stderr[:300]}")
return {
"ok": False,
"message": f"Image build failed: {build_result.stderr[:300]}",
"backup": backup_path,
}
except subprocess.TimeoutExpired:
_set_status("failed", "build", "Image build timed out after 600s.")
return {"ok": False, "message": "Image build timed out after 600s.", "backup": backup_path}
logger.info("Phase A complete — image built successfully.")
_set_status("running", "restart", "Container wird neu gestartet …")
# Phase B — swap the container using a helper container.
# When compose recreates our container, ALL processes inside die (PID namespace
@@ -388,6 +424,7 @@ def trigger_update(config: Any, db_path: str) -> dict:
raise ValueError("Could not find /app-source mount")
except Exception as exc:
logger.error("Failed to discover host source path: %s", exc)
_set_status("failed", "restart", f"Could not find host source path: {exc}")
return {"ok": False, "message": f"Could not find host source path: {exc}", "backup": backup_path}
logger.info("Host source directory: %s", host_source_dir)
@@ -426,6 +463,10 @@ def trigger_update(config: Any, db_path: str) -> dict:
)
if result.returncode != 0:
logger.error("Failed to start updater container: %s", result.stderr.strip())
_set_status(
"failed", "restart",
f"Update-Container konnte nicht gestartet werden: {result.stderr.strip()[:200]}",
)
return {
"ok": False,
"message": f"Update-Container konnte nicht gestartet werden: {result.stderr.strip()[:200]}",
@@ -434,6 +475,7 @@ def trigger_update(config: Any, db_path: str) -> dict:
logger.info("Phase B: updater container started — this container will restart in ~5s.")
except Exception as exc:
logger.error("Failed to launch updater: %s", exc)
_set_status("failed", "restart", f"Updater launch failed: {exc}")
return {"ok": False, "message": f"Updater launch failed: {exc}", "backup": backup_path}
return {
+20
View File
@@ -1,5 +1,25 @@
/* NetBird MSP Appliance - Custom Styles */
/* Sortable table headers */
.sortable-th {
cursor: pointer;
user-select: none;
white-space: nowrap;
}
.sortable-th:hover {
color: var(--bs-primary);
}
.sortable-th .sort-icon {
opacity: 0.35;
}
.sortable-th.sort-asc .sort-icon,
.sortable-th.sort-desc .sort-icon {
opacity: 1;
}
/* i18n FOUC prevention */
body.i18n-loading #login-page,
body.i18n-loading #app-page {
+6 -6
View File
@@ -254,13 +254,13 @@
<table class="table table-hover mb-0">
<thead class="table-light">
<tr>
<th data-i18n="dashboard.thId">ID</th>
<th data-i18n="dashboard.thName">Name</th>
<th data-i18n="dashboard.thSubdomain">Subdomain</th>
<th data-i18n="dashboard.thStatus">Status</th>
<th class="sortable-th" data-sort-col="id" onclick="setCustomerSort('id')"><span data-i18n="dashboard.thId">ID</span><i class="bi bi-arrow-down-up sort-icon ms-1"></i></th>
<th class="sortable-th" data-sort-col="name" onclick="setCustomerSort('name')"><span data-i18n="dashboard.thName">Name</span><i class="bi bi-arrow-down-up sort-icon ms-1"></i></th>
<th class="sortable-th" data-sort-col="subdomain" onclick="setCustomerSort('subdomain')"><span data-i18n="dashboard.thSubdomain">Subdomain</span><i class="bi bi-arrow-down-up sort-icon ms-1"></i></th>
<th class="sortable-th" data-sort-col="status" onclick="setCustomerSort('status')"><span data-i18n="dashboard.thStatus">Status</span><i class="bi bi-arrow-down-up sort-icon ms-1"></i></th>
<th data-i18n="dashboard.thDashboard">Dashboard</th>
<th data-i18n="dashboard.thDevices">Devices</th>
<th data-i18n="dashboard.thCreated">Created</th>
<th class="sortable-th" data-sort-col="max_devices" onclick="setCustomerSort('max_devices')"><span data-i18n="dashboard.thDevices">Devices</span><i class="bi bi-arrow-down-up sort-icon ms-1"></i></th>
<th class="sortable-th" data-sort-col="created_at" onclick="setCustomerSort('created_at')"><span data-i18n="dashboard.thCreated">Created</span><i class="bi bi-arrow-down-up sort-icon ms-1"></i></th>
<th data-i18n="dashboard.thActions">Actions</th>
</tr>
</thead>
+102 -9
View File
@@ -12,6 +12,8 @@ let currentPage = 'dashboard';
let currentCustomerId = null;
let currentCustomerData = null;
let customersPage = 1;
let customersSortBy = 'id';
let customersSortOrder = 'asc';
let brandingData = { branding_name: 'NetBird MSP Appliance', branding_logo_path: null, version: 'alpha-1.1' };
let azureConfig = { azure_enabled: false };
@@ -458,7 +460,7 @@ async function loadStats() {
async function loadCustomers() {
const search = document.getElementById('search-input').value;
const status = document.getElementById('status-filter').value;
let url = `/customers?page=${customersPage}&per_page=25`;
let url = `/customers?page=${customersPage}&per_page=25&sort_by=${customersSortBy}&sort_order=${customersSortOrder}`;
if (search) url += `&search=${encodeURIComponent(search)}`;
if (status) url += `&status=${encodeURIComponent(status)}`;
@@ -470,7 +472,32 @@ async function loadCustomers() {
}
}
function setCustomerSort(column) {
if (customersSortBy === column) {
customersSortOrder = customersSortOrder === 'asc' ? 'desc' : 'asc';
} else {
customersSortBy = column;
customersSortOrder = 'asc';
}
customersPage = 1;
loadCustomers();
}
function updateSortHeaders() {
document.querySelectorAll('.sortable-th').forEach(th => {
const col = th.getAttribute('data-sort-col');
const icon = th.querySelector('.sort-icon');
th.classList.remove('sort-asc', 'sort-desc');
if (icon) icon.className = 'bi bi-arrow-down-up sort-icon ms-1';
if (col === customersSortBy) {
th.classList.add(customersSortOrder === 'asc' ? 'sort-asc' : 'sort-desc');
if (icon) icon.className = `bi bi-arrow-${customersSortOrder === 'asc' ? 'up' : 'down'} sort-icon ms-1`;
}
});
}
function renderCustomersTable(data) {
updateSortHeaders();
const tbody = document.getElementById('customers-table-body');
if (!data.items || data.items.length === 0) {
tbody.innerHTML = `<tr><td colspan="8" class="text-center text-muted py-4">${t('dashboard.noCustomers')}</td></tr>`;
@@ -669,9 +696,9 @@ async function confirmDeleteCustomer() {
// ---------------------------------------------------------------------------
// Customer Actions (start/stop/restart/deploy)
// ---------------------------------------------------------------------------
async function customerAction(id, action) {
async function customerAction(id, action, name) {
if (action === 'deploy') {
showRedeployModal(id);
showRedeployModal(id, name);
return;
}
try {
@@ -683,9 +710,12 @@ async function customerAction(id, action) {
}
}
function showRedeployModal(id) {
const row = document.querySelector(`tr[data-customer-id="${id}"]`);
const name = row ? row.querySelector('td')?.textContent?.trim() : `#${id}`;
function showRedeployModal(id, name) {
// Prefer passed name, fallback to dashboard table row, then ID
if (!name) {
const row = document.querySelector(`tr[data-customer-id="${id}"]`);
name = row ? row.querySelector('td')?.textContent?.trim() : `#${id}`;
}
document.getElementById('redeploy-customer-id').value = id;
document.getElementById('redeploy-customer-name').textContent = name;
new bootstrap.Modal(document.getElementById('redeploy-modal')).show();
@@ -786,7 +816,7 @@ async function viewCustomer(id) {
<button class="btn btn-success btn-sm me-1" onclick="customerAction(${id},'start')"><i class="bi bi-play-circle me-1"></i>${t('customer.start')}</button>
<button class="btn btn-warning btn-sm me-1" onclick="customerAction(${id},'stop')"><i class="bi bi-stop-circle me-1"></i>${t('customer.stop')}</button>
<button class="btn btn-info btn-sm me-1" onclick="customerAction(${id},'restart')"><i class="bi bi-arrow-repeat me-1"></i>${t('customer.restart')}</button>
<button class="btn btn-outline-primary btn-sm me-1" onclick="customerAction(${id},'deploy')"><i class="bi bi-rocket me-1"></i>${t('customer.reDeploy')}</button>
<button class="btn btn-outline-primary btn-sm me-1" data-customer-name="${esc(data.name)}" onclick="customerAction(${id},'deploy',this.dataset.customerName)"><i class="bi bi-rocket me-1"></i>${t('customer.reDeploy')}</button>
<button class="btn btn-outline-warning btn-sm" id="btn-update-images-detail" onclick="updateCustomerImagesFromDetail(${id})">
<span id="update-detail-spinner" class="spinner-border spinner-border-sm d-none me-1"></span>
<i class="bi bi-arrow-repeat me-1"></i>${t('customer.updateImages')}
@@ -1371,11 +1401,12 @@ async function loadVersionInfo() {
if (needsUpdate) {
html += `<div class="mt-3">
<button class="btn btn-warning" onclick="triggerUpdate()">
<button class="btn btn-warning" id="update-trigger-btn" onclick="triggerUpdate()">
<span class="spinner-border spinner-border-sm d-none me-1" id="update-spinner"></span>
<i class="bi bi-arrow-repeat me-1"></i>${t('settings.triggerUpdate')}
</button>
<div class="text-muted small mt-1">${t('settings.updateWarning')}</div>
<div class="small mt-2 d-none" id="update-progress-text"></div>
</div>`;
}
el.innerHTML = html;
@@ -1387,14 +1418,76 @@ async function loadVersionInfo() {
async function triggerUpdate() {
if (!confirm(t('settings.confirmUpdate'))) return;
const spinner = document.getElementById('update-spinner');
const btn = document.getElementById('update-trigger-btn');
const progressText = document.getElementById('update-progress-text');
const setProgress = (msg) => {
if (!progressText) return;
progressText.classList.remove('d-none');
progressText.textContent = msg;
};
const stopUpdateUi = () => {
if (spinner) spinner.classList.add('d-none');
if (btn) btn.disabled = false;
};
if (spinner) spinner.classList.remove('d-none');
if (btn) btn.disabled = true;
setProgress(t('settings.updateStepStarting'));
try {
const data = await api('POST', '/settings/update');
showSettingsAlert('success', data.message || t('messages.updateStarted'));
} catch (err) {
showSettingsAlert('danger', t('errors.failed', { error: err.message }));
if (spinner) spinner.classList.add('d-none');
stopUpdateUi();
return;
}
// Phase 1: poll build/pull progress until the container restarts
// (connection drops, which is expected and is our cue to move to phase 2).
const stepLabelKey = {
backup: 'settings.updateStepBackup',
pull: 'settings.updateStepPull',
build: 'settings.updateStepBuild',
restart: 'settings.updateStepRestart',
};
for (let i = 0; i < 200; i++) {
await new Promise(r => setTimeout(r, 2000));
try {
const st = await api('GET', '/settings/update/status');
if (st.state === 'failed') {
showSettingsAlert('danger', st.message || t('errors.requestFailed'));
stopUpdateUi();
return;
}
setProgress(t(stepLabelKey[st.step] || 'settings.updateStepStarting') + (st.message ? `${st.message}` : ''));
} catch (err) {
// Connection dropped — the container is very likely mid-restart. Move on.
break;
}
}
// Phase 2: wait for the app to come back up, then reload version info.
setProgress(t('settings.updateStepReconnecting'));
for (let i = 0; i < 90; i++) {
await new Promise(r => setTimeout(r, 2000));
try {
await api('GET', '/settings/version');
setProgress(t('settings.updateStepDone'));
stopUpdateUi();
showSettingsAlert('success', t('settings.updateStepDone'));
await loadVersionInfo();
return;
} catch (err) {
// still restarting — keep polling
}
}
// Gave up waiting — surface this instead of spinning forever.
stopUpdateUi();
setProgress('');
if (progressText) progressText.classList.add('d-none');
showSettingsAlert('warning', t('settings.updateStepTimeout'));
}
// ---------------------------------------------------------------------------
+8
View File
@@ -230,6 +230,14 @@
"triggerUpdate": "Update starten",
"updateWarning": "Die App ist während des Rebuilds ca. 60 Sekunden nicht verfügbar.",
"confirmUpdate": "Update jetzt starten? Die Datenbank wird zuerst gesichert. Die App startet neu (~60 Sekunden Ausfallzeit).",
"updateStepStarting": "Update wird gestartet …",
"updateStepBackup": "Datenbank wird gesichert …",
"updateStepPull": "Code wird geholt …",
"updateStepBuild": "Docker-Image wird gebaut (kann mehrere Minuten dauern) …",
"updateStepRestart": "Container wird neu gestartet …",
"updateStepReconnecting": "Container startet neu — warte auf Verbindung …",
"updateStepDone": "Update abgeschlossen.",
"updateStepTimeout": "Update läuft länger als erwartet. Bitte Server-Logs prüfen oder die Seite in ein paar Minuten neu laden.",
"gitTitle": "Git-Repository Einstellungen",
"gitRepoUrl": "Repository URL",
"gitRepoUrlHint": "Wird für Versionsprüfungen und One-Click-Updates via Gitea API verwendet.",
+8
View File
@@ -262,6 +262,14 @@
"triggerUpdate": "Start Update",
"updateWarning": "The app will be unavailable for ~60 seconds during rebuild.",
"confirmUpdate": "Start the update now? The database will be backed up first. The app will restart (~60 seconds downtime).",
"updateStepStarting": "Starting update …",
"updateStepBackup": "Backing up database …",
"updateStepPull": "Fetching code …",
"updateStepBuild": "Building Docker image (can take several minutes) …",
"updateStepRestart": "Restarting container …",
"updateStepReconnecting": "Container is restarting — waiting for connection …",
"updateStepDone": "Update complete.",
"updateStepTimeout": "Update is taking longer than expected. Check the server logs or reload this page in a few minutes.",
"gitTitle": "Git Repository Settings",
"gitRepoUrl": "Repository URL",
"gitRepoUrlHint": "Used for version checks and one-click updates via Gitea API.",