The internet slows down at 11 a.m. Video calls freeze, a cloud accounting page stops loading and someone restarts the router. The office works again for twenty minutes, then the complaint returns. A bigger internet plan is proposed before anyone has checked which part of the network is failing.
Office network downtime can originate in a laptop, a wireless cell, a cable, a switch, a firewall, a power supply, an ISP or a cloud application. Several faults produce similar symptoms. The useful first question is which users and services are affected, and what still works.
This guide gives office owners, facilities teams and IT administrators a practical method to diagnose slow internet, isolate recurring failures and define a recovery plan. It also explains what to measure and what a professional network handover should contain.
Downtime, slow throughput and poor call quality are different problems
| Problem | What users notice | Measurements to collect |
|---|---|---|
| Complete outage | No usable access to required services | Power state, links, reachability and event logs |
| Low throughput | Large transfers take longer than expected | Actual transfer rate, negotiated speed and utilisation |
| High latency | Pages and interactive actions feel delayed | Round-trip time to relevant destinations |
| Jitter | Audio arrives unevenly or calls sound broken | Variation in packet delay and application telemetry |
| Packet loss | Calls freeze or transfers retry | Loss at destination, interface errors and application reports |
| Service-specific failure | One website or application fails while others work | DNS, authentication, policies and provider status |
A speed test reports only one type of performance against a particular test endpoint at a particular moment. It cannot prove that office WiFi, a file server, every cloud application or a VPN is healthy. Keep a normal baseline so later measurements have context.
Start with the scope: one user, one area or the entire site
Ask for the failure time, the device, the connection method and the exact service. “Internet down” might mean an application login failed while all other traffic still works. Record observations before making changes.
- One endpoint: compare another device on the same connection and the affected device on a known-good connection.
- WiFi users only: compare a wired client using the same internet gateway.
- One floor or room: inspect its access switch, APs, power and uplink.
- One service: test other destinations and check application or DNS behaviour.
- All users: examine shared equipment, power, routing, DNS and WAN connectivity.
Draw the actual path from the user to the service
For a WiFi laptop, the path may be laptop → AP → access switch → core → firewall → ONT → ISP → cloud application. A local printer or NVR may use only part of that path. A VPN adds a tunnel and another destination network.
Mark which devices are common to the failed users. If a camera, AP and door controller all restart together on one floor, a shared PoE switch or power source deserves attention before three endpoint replacements.
1. LAN cabling, patch cords and termination faults
A damaged patch lead, poor termination or marginal permanent link can create intermittent errors. The connection may negotiate a lower rate, flap between up and down, or retry traffic while still showing a link light.
Start with the negotiated link speed and switch interface counters. Look for new errors, drops or link-state changes during the complaint. Cumulative counters from months ago are less useful than a time-correlated increase. Errors suggest a link problem but do not uniquely identify the cable.
Compare using a known-good patch lead and an appropriate alternate port, preserving the original settings. Check the NIC and switch configuration together. Do not force speed or duplex at one end while leaving an incompatible setting at the other.
For installed Cat6/Cat6A links, use appropriate certification against the project limit and match reports to port IDs. Certification measures the fixed cabling; operational checks measure the live network. Our Fluke Testing guide explains why continuity alone is insufficient.
2. WiFi coverage, contention and roaming
If a wired client works while nearby wireless clients struggle, focus on the radio path and AP uplink. Check which AP and band the affected client actually uses, signal quality, retransmissions, channel utilisation and the negotiated uplink rate.
Full signal bars do not prove capacity. A strong signal can share airtime with many active clients or neighbouring networks. More APs or maximum transmit power can also increase interference when placement and channels are poorly planned.
An AP above a metal ceiling, behind an RCC beam or in a cupboard may fail to cover the desks it serves. Movement between APs introduces another test: does the call remain usable when the client roams? Check the affected device model rather than assuming every client roams identically.
Validate at working positions and during representative use. Our WiFi access point placement guide covers survey, location and cable planning. A replacement AP should follow the findings, rather than merely a higher advertised speed.
3. Congestion and queueing: why a fast line can feel slow
Cloud backups, software updates, large uploads and guest downloads can compete with calls. Upload capacity is often overlooked. When a queue fills, interactive traffic may wait even though the connection still transfers data at a high rate.
Measure responsiveness both at idle and while representative traffic is running. Test latency, loss and application behaviour against the same destination. Correlate the result with WAN and uplink utilisation.
An appropriate response may include scheduling large transfers, controlling guest use or configuring supported traffic shaping and queue management. QoS cannot manufacture capacity and cannot control every bottleneck beyond the office. Validate any adjustment with the actual business applications.
Avoid assuming that every guest network slows staff traffic simply because it lacks a VLAN. VLANs separate traffic logically; shared radio airtime, uplinks and WAN capacity still require their own planning.
4. Switch uplinks, loops and shared infrastructure
A floor can have gigabit desk ports while its uplink runs at 100 Mbps. Inspect the real negotiated rates and utilisation across the complete path. Add office, camera and AP traffic that crosses the shared link.
An unintended Layer 2 loop can flood traffic and disrupt a wider area. Record the time of recent patching, MAC movement or topology events. Identify the suspected loop from the documented topology and switch evidence rather than disconnecting unrelated links.
Use supported spanning-tree and edge protections appropriate to the design. Configure redundant links as a proper topology or supported aggregation. Two parallel cables installed casually are not a tested redundancy plan.
Switch restarts can also follow supply failure, high temperature, firmware faults or exhausted resources. Check uptime, logs and environment. A switch becoming hot does not by itself establish the cause; relate temperature and events to the vendor operating limits.
5. DHCP, IP conflicts and DNS
DHCP gives a client its IP address, subnet, gateway and usually DNS settings. An exhausted scope, unreachable server or rogue DHCP source can affect new clients while existing clients keep working. Check the actual lease and options on an affected endpoint.
Duplicate static addresses can produce intermittent reachability or a device appearing under the wrong identity. Reconcile static assignments, reservations and the DHCP range. Record device MAC addresses and use the network evidence to locate conflicts.
DNS converts a name into an address. If network reachability works but name resolution fails, investigate the configured resolver, its reachability and any filtering. Compare known destinations and record DNS responses.
Browsing directly to a website’s IP address is not a reliable DNS test: HTTPS certificates, virtual hosting and application routing may require the hostname. Use a name-resolution query and then test the actual service with the correct name.
| Observation | Likely area to investigate | What it does not prove |
|---|---|---|
| Client has no valid expected lease | DHCP path, scope or local connectivity | That the ISP is down |
| Gateway reachable, public destinations fail | WAN, routing or policy | That the firewall must be replaced |
| Names fail to resolve; approved IP test works | DNS configuration or resolver path | That opening a website by IP will work |
| Same IP appears associated with different devices | Address conflict or stale records | Which endpoint is correct without verification |
6. Firewall policies, inspection and VPN paths
A firewall may restrict a required application, reach its inspected capacity or send traffic through an unsuitable path. Gather the affected source, destination, time and service, then check matching policies, session logs and resource counters.
Size using the features actually enabled and the business workload. Raw forwarding throughput is not the same as performance with inspection, logging and remote access. A published number still needs interpretation against the product test conditions.
Compare VPN and non-VPN paths only where the organisation permits it. A full-tunnel client may send cloud traffic through another office before reaching the service. That changes latency and capacity requirements.
Do not disable the entire security stack to make a speed test look better. Use targeted, controlled diagnostics with an administrator and verify the intended policy afterwards. If a change causes the issue, a tested rollback is often safer than multiple improvised edits.
Our office firewall setup guide covers zones, remote access and configuration ownership.
7. CCTV traffic: understand where the video travels
Local CCTV recording does not automatically consume internet bandwidth. A camera streaming to a local NVR generates LAN traffic; cloud recording or remote viewing can also use the WAN.
Draw recorder and viewer locations. If cameras and NVR connect to the same access switch, recording can stay within that switch. If the NVR is at the core, those streams cross the uplink. Viewing from other floors and retrieving clips adds traffic along its own path.
For illustration, sixteen cameras averaging 8 Mbps produce 128 Mbps of video payload before overhead, extra streams or other traffic. A 100 Mbps shared uplink cannot carry that load. This says nothing by itself about the ISP plan, because the recording may remain on the LAN.
A CCTV VLAN can improve organisation and access control, but does not create additional physical bandwidth. Check camera bitrates, recorder capacity, uplinks and permitted viewers together.
8. PoE and power failures
An AP disappearing can be a power fault rather than a wireless fault. A camera failing at night can coincide with infrared or heater demand. Check per-port allocations, the total PoE budget and event logs under the relevant load.
Use the exact endpoint requirements, compatible PoE type and actual switch power configuration. Our PoE switch guide provides a worked sizing example and explains why spare sockets do not guarantee spare power.
Include the ONT, firewall, core switches and required access switches in the backup scope. Keeping one server alive does not preserve its network path when the switch loses power.
Verify UPS runtime at representative load and monitor battery condition. Record device uptime to correlate brief power interruptions with observed outages. Scheduled recovery tests should confirm configuration persistence and orderly restart.
9. ISP and external service failures
When local services remain healthy but external traffic fails, inspect WAN link state, addressing, default routes and firewall events. Record ONT indicators and provider information according to the supported troubleshooting procedure.
A second provider or permitted mobile connection can help compare the destination. If an application fails across independent connections while other services work, investigate that application’s status, account or region. The comparison narrows the fault rather than proving a provider is blameless.
Send the ISP a useful incident: timestamps with timezone, affected circuit, relevant destination tests, packet-loss observations and equipment status. Repeated screenshots of a single speed test rarely explain an intermittent path failure.
A controlled troubleshooting sequence
- Record the symptom: exact time, affected users, access method and failed service.
- Compare a known-good endpoint and a wired client where appropriate.
- Check local lease, gateway, DNS and the permitted application path.
- Inspect cable/link state, interface counters and the shared uplink.
- Correlate AP, switch, firewall, UPS and WAN logs at the failure time.
- Run repeatable performance tests at idle and under representative load.
- Apply the smallest evidence-based fix, with a configuration backup and rollback method.
- Retest the original complaint and monitor for recurrence.
Useful Windows checks—and how to interpret them
The following examples are read-only diagnostics. Replace example addresses and names with approved site targets. The sample gateway is not a claim about your office IP plan.
| Check | Example command | Interpretation |
|---|---|---|
| IP configuration | ipconfig /all | Inspect adapter, address, gateway, DHCP and DNS details |
| Gateway reachability | ping -n 30 10.0.10.1 | Compare local response during the issue |
| Name resolution | nslookup example.com | Check the resolver and returned records |
| Route observation | tracert example.com | Observe the responding route hops |
| HTTPS TCP reachability | Test-NetConnection example.com -Port 443 | PowerShell test for TCP connection to that destination |
| Adapter state | Get-NetAdapter | PowerShell view of adapter status and reported link speed |
A ping failure can result from filtering, and a successful ping does not prove video or HTTPS works. Intermediate traceroute hops may rate-limit diagnostic replies while forwarding ordinary traffic normally. Loss shown at one hop, without corresponding downstream loss, is not enough to identify a failing link.
TCP connection success does not prove application login, TLS operation or acceptable performance. Use these tools alongside the application’s own diagnostics and the network logs.
Worked incident: every video call fails when backup starts
Illustrative scenario: a 25-person office reports afternoon call freezes. Wired and wireless users are affected. Local file access remains normal. The primary WAN upload becomes heavily utilised at the same time as a scheduled off-site backup.
The investigation compares latency at idle and during backup, checks firewall traffic and application call telemetry, and finds the symptom follows the busy upload period. This evidence points towards the shared WAN path; it does not justify replacing every AP.
The team schedules bulk transfer outside the busy period and applies an appropriate supported traffic-management configuration. It then repeats a real call while representative background traffic runs. The incident closes only after that validation and a monitoring period.
A different result would require a different fix. If only one meeting room failed while the wired WAN test remained healthy, AP airtime, placement and the room’s clients would be higher priorities.
Backup internet and failover: define the promise
Dual WAN can reduce the effect of one provider outage. It does not protect a failed shared firewall, core switch, UPS or cable route. Check whether the providers share a last-mile path or upstream dependencies.
Health checks should test useful connectivity, not merely whether an Ethernet port is up. Agree which applications the backup line must support and what traffic is restricted when it is active.
Changing the public IP can interrupt active calls, VPNs and authenticated sessions. Failover may restore new connections while existing sessions need to reconnect. Test both failure and return to the primary link so recovery does not repeatedly flap.
For critical sites, additional resilience may include a supported firewall pair, redundant core design and separate power arrangements. The appropriate level follows the business impact and tested operating procedure.
Measure availability and recovery rather than promising zero downtime
Define the service and the measurement period. “Internet available” and “business application usable” are not identical. Record incidents, duration, affected users and whether the time was planned maintenance.
For an illustrative 22-day month with eight working hours per day, the service window is 176 hours. Two hours of unplanned outage in that window gives about 98.86% availability. A different reporting window changes the result, so state the denominator.
Agree who receives alerts and who can act. Monitoring should distinguish a powered-off client from a failed business service where possible. Track recovery duration and repeat faults so the team addresses causes rather than counting reboots.
Preventive maintenance and monitoring
- Monitor shared uplink utilisation, new interface errors and unexpected link changes.
- Track PoE allocation, device restarts, cabinet conditions and UPS health.
- Use AP and application telemetry to review recurring wireless or call issues.
- Keep configuration backups and a documented, supported firmware update process.
- Review DHCP capacity, static address records and DNS dependencies.
- Inspect labels, patching and rack ventilation after changes.
- Schedule backup traffic and validate critical applications during busy periods.
- Test failover, restoration and operator escalation at agreed intervals.
Synchronise equipment clocks where supported and record timezone consistently. Logs from an AP, switch and firewall are far more useful when the same incident can be aligned across them.
What a network service handover should contain
| Deliverable | Why it reduces downtime |
|---|---|
| Topology and rack elevation | Identifies shared equipment and dependency paths |
| Port map and cable reports | Connects device locations with tested installed links |
| IP, VLAN and DHCP schedule | Reduces address conflicts and configuration guesswork |
| AP layout and survey findings | Provides a baseline for coverage and capacity |
| Firewall and switch backups | Supports controlled rollback and replacement |
| Power / UPS and failover results | Defines which services survive each tested failure |
| ISP details and escalation contacts | Makes provider incidents actionable |
| Maintenance ownership and secure access | Ensures the right person can diagnose and recover |
Credentials belong with authorised owners through a secure handover process. Device maps and backups should remain available even when the main office network is unavailable.
Frequently asked questions
Should we upgrade the internet plan first?
Upgrade when measured demand and the WAN capacity show a real shortfall. A faster ISP plan will not repair poor WiFi placement, a 100 Mbps internal bottleneck, failed DNS or an unstable cable.
Why does restarting the router help temporarily?
A restart can clear transient state or re-establish a connection, but it can also erase useful evidence. Repeated improvement is a symptom to investigate with logs and uptime, rather than proof that a replacement is required.
Can guest WiFi and CCTV cause office slowdown?
They can compete on shared links, radios or internet paths. Find the traffic path and utilisation first. Separate VLANs help policy and organisation, while bandwidth planning addresses capacity.
Does packet loss on traceroute prove the ISP has a bad router?
No. Some intermediate devices limit diagnostic replies. Compare the destination result, later hops and real application performance before assigning the fault.
Will dual internet keep every call uninterrupted?
Not necessarily. A WAN change can require session reconnection. Define the recovery target and test the actual calling, VPN and business applications.
APYS Projects: diagnose the complete network path
APYS Projects handles office networking, structured cabling, Fluke Testing, WiFi planning, firewall setup, PoE switching, CCTV and ELV integration. We can inspect the installed path, identify shared dependencies and document corrective work with relevant tests.
A useful investigation starts with your topology, affected locations, device details and failure timestamps. The result should explain the cause, the change made and how the original symptom was retested.
Facing repeated office network downtime? Share the symptoms and site details with Purchase@apysprojects.com or +91 9921490342. Contact APYS Projects for network assessment, corrective work and documented handover.
Technical references
Further guidance: Cisco TCP/IP troubleshooting, Cisco interface and NIC compatibility troubleshooting, Microsoft Teams meeting diagnostics, Microsoft Teams network preparation and Cloudflare diagnostic route interpretation. Product-specific commands, thresholds and changes should follow the installed platform documentation and the site requirements.