Cloud System Engineer

Tungsten Automation
Tungsten Automation

Other Engineering, IT

Sofia City Province, Bulgaria

EUR 29k-34k / year

Posted on Jul 30, 2026

Cloud System Engineer



Tracking Code

E26-107

Job Location

Business Centre "Labirint" 5th Floor Liulin 10 District, Sofia, Sofia,

Job Level

Not Applicable

Category

Cloud

Position Type

Full-Time/Regular

Job Purpose

The Cloud Systems Engineer – Application Reliability & Investigations is a diagnostic engineering role responsible for deep-dive troubleshooting, root cause analysis, and resolution of complex issues across the enterprise application ecosystem.

As a key member of the Investigations Team, you will diagnose failures involving internal system integrations, B2B communications, and application performance anomalies. You will work closely with Support, Development, Security, and Cloud Operations teams to identify root causes, resolve incidents, and implement long-term improvements.

While collaborating with Cloud Operations on infrastructure and networking matters, your primary focus will be on application-layer diagnostics, transaction flow analysis, and maintaining end-to-end system reliability. This role provides an excellent opportunity to grow technical expertise under the guidance of senior engineers while progressively taking ownership of increasingly complex investigations.

Key Responsibilities

1. Technical Investigations (Primary Focus)

  • Conduct detailed investigations into escalated incidents involving multiple applications, systems, and integration points.
  • Analyze transaction failures across APIs, databases, message queues, and third-party integrations.
  • Parse and analyses structured data formats including XML, JSON, EDI, and flat files to identify data-related issues.
  • Correlate application logs, traces, and performance metrics using observability platforms such as Splunk, Logz.io and New Relic to identify root causes.
  • Investigate integration failures between internal platforms, identifying configuration issues, orchestration failures, and data inconsistencies.
  • Produce clear and comprehensive investigation reports to support knowledge sharing and long-term problem prevention.

2. Incident Response & Collaboration

  • Respond to Priority 1–5 incidents by performing rapid triage and technical assessments.
  • Provide investigation support throughout the lifecycle of critical (P1/P2) incidents.
  • Participate in post-incident reviews and contribute root cause analysis and preventive recommendations.
  • Communicate technical findings effectively to both technical and business stakeholders.
  • Maintain awareness of ongoing incidents and proactively identify related issues across dependent systems.

3. Application & Integration Troubleshooting

  • Troubleshoot application communication protocols including AS2, SFTP, and HTTPS.
  • Diagnose authentication and authorization failures in collaboration with Security and Cloud Operations teams.
  • Investigate data transformation, mapping, and processing errors across internal systems and external partners.
  • Support Cloud Operations by providing application-level analysis during infrastructure or network investigations.

4. Observability & Monitoring

  • Enhance monitoring capabilities using New Relic, Splunk, Logz.io, and Incident.io.
  • Develop dashboards and alerts to monitor transaction volumes, success rates, and application response times.
  • Identify monitoring gaps during investigations and recommend improvements to telemetry and observability.
  • Collaborate with Development and Cloud Operations teams to improve logging and monitoring standards.

5. Automation & Continuous Improvement

  • Develop automation scripts using Python, PowerShell, or Bash to streamline diagnostics and investigations.
  • Create tools for automated log collection, health checks, and data validation.
  • Identify recurring issues and recommend preventive improvements through configuration, code, or process enhancements.
  • Maintain runbooks, troubleshooting guides, and technical knowledge base documentation.
  • Develop automated integration validation to reduce regression risks.
  • Utilize AI-enabled tools (e.g., chatbots, documentation automation, analytics assistants) to improve efficiency, accuracy, and streamline routine tasks while following company AI governance and data privacy standards

6. Internal Systems & Product Knowledge

  • Develop comprehensive knowledge of internal products, system architecture, and integration flows.
  • Partner with Product and Development teams to support operational readiness for new releases.
  • Document end-to-end transaction flows, dependencies, and potential failure points.
  • Participate in integration testing for system upgrades and partner onboarding activities.

7. Continuous Learning & Knowledge Sharing

  • Identify opportunities to improve investigation processes and tooling.
  • Share technical knowledge and best practices across Engineering and Support teams.
  • Stay informed on emerging technologies, integration standards, and industry best practices.

Base salary range: For this role, based in Bulgaria, the base salary range is €29,000-34,0000 per year, set using objective, gender-neutral job evaluation criteria applied to all employees performing this role or work of equal value.

Determining your actual pay: Your base salary within this range will be based on objective, gender-neutral factors — work location, skills, qualifications, experience, and education/training. It will not be reduced based on prior salary or negotiating position, nor will it fall below the range or below what colleagues in equal work receive.

Scope: This range covers base salary only, not bonuses, benefits, or other remuneration (as applicable).

Your rights: Tungsten Automation does not ask about current or past pay, and you need not disclose it. This notice does not limit your right to negotiate or your entitlement to equal pay for equal work or work of equal value.

While the job description describes what is anticipated as the requirements of the position, the job requirements are subject to change based upon any changing needs and requirements of the business.


Required Skills

Core Competencies

Technical Excellence

  • Application troubleshooting
  • Root cause analysis
  • Systems integration
  • Observability and monitoring
  • Automation and scripting

Communication

  • Incident communication
  • Technical documentation
  • Stakeholder management
  • Cross-functional collaboration

Problem Solving

  • Analytical thinking
  • Structured investigation
  • Decision making
  • Continuous improvement

Personal Attributes

  • Analytical Mindset – Able to break down complex problems and identify root causes through systematic investigation.
  • Intellectual Curiosity – Passionate about understanding systems and solving challenging technical problems.
  • Persistence – Demonstrates resilience when investigating complex or intermittent issues.
  • Calm Under Pressure – Maintains composure during critical incidents and communicates effectively.
  • Collaborative – Works effectively across teams and shares knowledge openly.
  • Growth-Oriented – Continuously develops technical knowledge and embraces feedback.
  • Process-Driven – Values documentation, repeatable processes, and operational excellence.
  • Customer-Focused – Understands the importance of reliability and service quality for customers and partners.

Required Experience

Required Qualifications

  • Bachelor's Degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
  • 1–3 years' experience in Application Support, Systems Engineering, Site Reliability Engineering (SRE), or similar technical support roles.
  • Experience troubleshooting distributed applications and complex system integrations.
  • Working knowledge of AS2, SFTP, HTTPS, and related communication protocols.
  • Experience working with structured data formats such as XML, JSON, EDI, and flat files.
  • Scripting experience using Python, PowerShell, or Bash.
  • Hands-on experience with monitoring and observability platforms such as New Relic, Splunk, Logz.io, or equivalent.
  • Basic knowledge of Linux and Windows application hosting environments (IIS, Apache, Tomcat).
  • Understanding of networking fundamentals including TCP/IP, DNS, Firewalls, and TLS/SSL.
  • Experience with relational and NoSQL databases including Oracle, MySQL, PostgreSQL, and MongoDB.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent written and verbal communication skills.
  • Skills in prompting AI systems and assessing output quality.
  • Ability to leverage AI to ideate, develop, and scale to the needs of their department.

Preferred Qualifications

  • Experience with Docker and Kubernetes.
  • Knowledge of CI/CD tools such as Jenkins and GitHub Actions.
  • Cloud certifications (AWS, Azure, or equivalent), or actively working towards certification.
  • Experience with EDI standards and B2B integration platforms.
  • Familiarity with messaging technologies including RabbitMQ, Kafka, or Amazon SQS.
  • Experience troubleshooting REST APIs and web services.

Tungsten Automation Corporation is an Equal Opportunity Employer M/F/Disability/Vets