QA/Test Engineer - Google Cloud Observability Stack

Posted 1 hour 55 minutes ago by ComTech Europe Limited

40,00 € Hourly
Contract
Not Specified
Other
Not Specified, Poland
Job Description

Description

My client are looking for a QA/Test Engineer to work as part of a team delivering a migration from AWS to Google Cloud.

About the role

We're looking for a QA/Test Engineer to be the independent voice of "does it actually work" on a large-scale cloud migration programme: ten public-facing applications moving from a fragmented estate (AWS, an in-house Kubernetes platform, on-premises VMs, and a SaaS CMS) onto Google Cloud. You won't be building the applications - you'll be proving, with measured evidence, that each one performs at least as well after migration as it did before, and that it can recover when something goes wrong. Your evidence is what customer sign-off and program milestone payments are built on, so it needs to hold up on its own - independent of the team that built what you're validating.

You'll be involved from day one, establishing what "good" looks like before anything moves, and you'll stay involved through the last cutover, making the call on whether a migrated service is genuinely ready for production traffic.

What you'll do

Establish pre-migration performance baselines: traffic patterns, cache hit ratios, latency, and current cost, drawn from load balancer, CDN, and analytics data over a real measurement window - not a guess.
Define validation criteria and test cases per application, covering both functional behaviour and non-functional targets (availability, latency, recovery time/point objectives).
Design and run load tests against each migrated public service at its measured peak request rate plus an agreed safety margin, and record throughput, 95th-percentile latency, error rate, and cache hit ratio as acceptance evidence.
Design and execute recovery time/recovery point objective tests for every application, and produce the evidence that demonstrates the target is actually met - not just documented.
Compile the acceptance evidence that feeds each migration phase's go/no-go decision and sign-off package.
Monitor a parallel-run cutover for the highest-risk service in the estate (a shared image delivery service under heavy load), watching error rate and latency against agreed thresholds and calling for rollback if they're breached.
Write the post-cutover review for each phase: what the metrics showed against baseline, what came up, and whether the new path is trusted enough to let the old one go.

What you'll bring

Real experience running exactly this kind of exercise before: measuring a system's behaviour before a migration or major change, and proving - with comparable, reproducible numbers - that it still holds up afterward. This matters more to us than which specific tool you used to do it.
Working fluency with a load-testing tool (k6, Locust, JMeter, Gatling, or equivalent) - we're tool-agnostic as long as you can justify the choice and get a fair, repeatable comparison out of it.
Comfort working with Google Cloud's observability stack (Cloud Monitoring, Cloud Trace) to both source the baseline and verify the post-migration result - our target architecture leans on Cloud Monitoring SLOs as the source of truth for latency and availability targets, so this isn't optional.
A methodical approach to recovery testing: you can design a disaster-recovery or failover drill, run it, and write it up in a way a customer's auditors would accept.
Enough Scripting ability to automate the boring parts of baseline collection and before/after comparison, rather than doing it by hand every time.
Judgment to prioritise: several of these validation efforts will overlap in time across different applications, and you'll need to sequence your own work sensibly rather than needing someone else to do it for you.

Nice to have

Experience working against SLO-based acceptance criteria rather than ad hoc pass/fail checks.
Experience testing WAF/edge-security rule efficacy (eg Cloud Armor or an equivalent), including rate-limit enforcement.
Comfort being the one who says "not yet" to a cutover when the numbers don't support it, even under schedule pressure.

Location: you can work remotely but must be based in the EU, so unfortuantely no UK candidates can be considered for this role.

Start: mid of October (or beginning of November)

Allocation: 80%, might vary a slightly depending on the week so could suit someone with 4 days per week availability.

Language: English