What Product Reliability Jobs are in India?

Showing 625 Product Reliability jobs in India

Senior Manager, Product Development Engineering (Memory Reliability , NAND,Failure Analysis, Root...

Bengaluru SanDisk

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

**Company Description**
Sandisk understands how people and businesses consume data and we relentlessly innovate to deliver solutions that enable today's needs and tomorrow's next big ideas. With a rich history of groundbreaking innovations in Flash and advanced memory technologies, our solutions have become the beating heart of the digital world we're living in and that we have the power to shape.
Sandisk meets people and businesses at the intersection of their aspirations and the moment, enabling them to keep moving and pushing possibility forward. We do this through the balance of our powerhouse manufacturing capabilities and our industry-leading portfolio of products that are recognized globally for innovation, performance and quality.
Sandisk has two facilities recognized by the World Economic Forum as part of the Global Lighthouse Network for advanced 4IR innovations. These facilities were also recognized as Sustainability Lighthouses for breakthroughs in efficient operations. With our global reach, we ensure the global supply chain has access to the Flash memory it needs to keep our world moving forward.
**Job Description**
We're looking for an experienced Senior Manager to lead our Memory Reliability & Qualification engineering organization in Bengaluru, India. In this strategic leadership role, you will oversee the complete NAND Flash memory qualification lifecycle, drive innovation, and build a high-performing engineering team. You will be responsible for translating business objectives into technical strategies while maintaining operational excellence and delivering world-class solutions.
+ Lead and mentor a team of NAND Technology Reliability engineers, fostering a culture of innovation, collaboration, and continuous improvement
+ The candidate should have a vast experience in semiconductor industry along with strong skills in Failure Analysis, Root Cause Analysis, Device Physics and programming
+ Develop and execute memory qualification strategies aligned with organizational goals and business demands
+ Oversee the complete qualification lifecycle, from requirements gathering through successful memory qualification, ensuring timely delivery and attention to details
+ Manage engineering budgets, resource allocation, and project timelines to optimize operational efficiency
+ Make critical technical decisions based on the performance and reliability checkpoint data to judge the memory technology
+ Collaborate with cross-functional teams including product engineering, test engineering and product management to ensure seamless delivery
+ Develop and implement performance metrics to track engineering team productivity and deliverables quality
+ Represent the team in executive meetings and strategic planning sessions
**Qualifications**
**Required Qualifications:**
+ MS Degree in Electrical Engineering, Applied Physics, or a related field.
+ 13+ years of experience in reliability testing, product development engineering or failure analysis
+ 7+ years of experience in a management role handling a team size of at least 10 members
+ Strong technical knowledge in semiconductor domain, device physics, data analysis/analytics
+ Demonstrated ability to lead, develop, and scale engineering teams
+ Excellent project management skills with experience managing complex, multi-disciplinary projects
+ Strong analytical and problem-solving abilities with attention to detail
+ Proficiency in budget management and resource planning
+ Excellent communication and stakeholder management skills
**Preferred Qualifications:**
+ Experience in NAND Flash Reliability testing is a big plus
+ Experience in semiconductor, automotive or electronics industries
+ Familiarity with Agile and Scrum methodologies
+ C Programming & Scripting (Python/Shell scripting).
+ Experience with cross-functional team leadership in matrix organizations
**Additional Information**
Sandisk thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.
Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.
Is this job a match or a miss?
Apply Now

Engineer - Reliability

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

**What you'll do:**
**Eaton announced, on January 26, 2026, the intent to separate its Mobility Group (including both the Vehicle and eMobility segments) into an independent, publicly traded company. We expect to complete the separation by the end of the first quarter of 2027. As a standalone company, the Mobility business will be more focused and agile, creating exciting opportunities for employees to grow, innovate, and help shape the future of mobility. The separation reflects strong confidence in the Mobility team and positions the new company as an employer of choice in the automotive and commercial vehicle industry.**
The compensation and benefits that will initially be offered for this position are based on Eaton's plans, programs and practices. If you are offered and accept this position and are actively employed by the Mobility Group when the spinoff closes, the new company will provide further details to employees concerning compensation and benefits at that time.
Responsibility:
- Create and/or revise DFR deliverables such as Reliability goal setting, Reliability Predictions, Failure Mode Effects Analysis (FMEA), Accelerated Life Test , Reliability Demonstration Test (RDT), Reliability Growth analysis ,FRACAS for NPIs, VAVE projects and/or existing product improvements.
- Work with cross functional team as well as global reliability engineering team for designing in reliability on NPI programs and/or existing product improvements,VAVE projects & responsible for identifying, managing and mitigating reliability risks under supervision of senior.
- Planning project milestones, executing milestones, conducting deliverables reviews, communication with stakeholder and on-job development through technical mentoring.
- Along with warranty monitoring, forecasting, reporting; develop actuarial models library for eMobility components
**Qualifications:**
Education level required : Bachelor's Degree in Electrical/Electronics Engineering. Preferred: Master's Degree in Reliability Engineering.
Experience : 0-1 Years
Technical knowledge :
- Skilled in the areas of reliability statistics, FMEA, FTA,8D,FRACAS, life modelling & analyzing field/warranty data into relevant life and failure rate information.
- Experience with any industry proven reliability analysis suite of tools like Minitab, Reliasoft preferred.
- Desired: Work experience in the Eletrical/Electronics systems preferably Power electronics(Inverter/Converters), Power systems, battery, inverters, protection circuits etc.
- Preferred: DFSS GB certified from reputed organization, Experience in safety/Reliablility assessment of E/E system
**Skills:**
- Process Management-Good at figuring out the processes necessary to get things done, knows how to organize people and activities, knows what to measure and how to measure it, Can simplify complex processes, Gets more out of fewer resources
- Problem Solving - Uses rigorous logic and methods to solve difficult problems with effective solutions, probes all fruitful sources for answers, can see hidden problems, Is excellent at honest analysis Looks beyond the obvious and doesn't stop at the first answers
- Decision quality - makes good decisions based upon a mixture of analysis, wisdom, experience, and judgment
- Drive for results - can be counted on to exceed goals successfully
- Interpersonal savvy - relates well to all kinds of people; builds appropriate rapport.
Is this job a match or a miss?
Apply Now

Senior Engineer, Reliability Engineering

Bengaluru SanDisk

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

**Company Description**
SanDisk understands how people and businesses consume data and we relentlessly innovate to deliver solutions that enable today's needs and tomorrow's next big ideas. With a rich history of groundbreaking innovations in Flash and advanced memory technologies, our solutions have become the beating heart of the digital world we're living in and that we have the power to shape.
SanDisk meets people and businesses at the intersection of their aspirations and the moment, enabling them to keep moving and pushing possibility forward. We do this through the balance of our powerhouse manufacturing capabilities and our industry-leading portfolio of products that are recognized globally for innovation, performance and quality.
SanDisk has two facilities recognized by the World Economic Forum as part of the Global Lighthouse Network for advanced 4IR innovations. These facilities were also recognized as Sustainability Lighthouses for breakthroughs in efficient operations. With our global reach, we ensure the global supply chain has access to the Flash memory it needs to keep our world moving forward.
**Job Description**
+ Excellent interpersonal skills and team player. Proven ability to work as part of a global team in multiple geographies
+ Must be able to work closely with the cross-functional team members on a daily basis, set up meetings and calls with the stake-holders and resolve issues diligently
+ Fostering a culture of innovation and continuous improvement
+ Recommend new technologies, tools, and methodologies to enhance our engineering capabilities
+ Able to methodically root cause complex failure mechanism
+ Understanding system specifications and memory requirements
+ Perform NAND flash storage based analysis and verification from product definition and planning through production release
+ Understanding of NAND flash system level behavior
+ Participates in cross functional meetings with memory, firmware and product line teams to develop flash storage products
+ Knowledge on the PLC process of a product
+ Provides feedbacks on new system features for next generation memory product development
+ Should be working with Factory teams, in co-ordinating on the items like drive builds Mechanical tests, etc.
+ Will be handling RCA, 8D and process improvements
+ Understanding of manufacturing operations and product/test engineering flow
+ Making sure product is meeting the reliability and quality spec.
**Qualifications**
+ Bachelor/Masters degree from Electrical/Electronics/Computer science preferred with 3+ years of experience in Reliability/Memory design/system engineering
+ Experience in memory design/characterization
+ Knowledge of Semiconductor physics and Computer Architecture
+ Knowledge on factory operations and Agile methodology is a plus
+ Knowledge of storage interfaces such as UFS, eMMC, SATA or PCIe is a plus
+ Experience in SSD/Embedded reliability and testing is a plus
+ Exceptional written and verbal communication skills
+ Proficient in Microsoft Office applications to prepare status report
**Additional Information**
Sandisk thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.
Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.
Is this job a match or a miss?
Apply Now

Site Reliablity Engineer III

Bengaluru Proofpoint

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

**About Us:**
Proofpoint is a global leader in human- and agent-centric cybersecurity. We protect how people, data, and AI agents connect across email, cloud, and collaboration tools. Over 80 of the Fortune 100, 10,000 large enterprises, and millions of smaller organizations trust Proofpoint to stop threats, prevent data loss, and build resilience across their people and AI workflows. Our mission is simple: safeguard the digital world and empower people to work securely and confidently. Join us in our pursuit to defend data and protect people.
**How We Work:**
At Proofpoint you'll be part of a global team that breaks barriers to redefine cybersecurity guided by our BRAVE core values:
**Bold** in how we dream and innovate
**Responsive** to feedback, challenges and opportunities
**Accountable** for results and best in class outcomes
**Visionary** in future focused problem-solving
**Exceptional** in execution and impact
Corporate Overview
Proofpoint is a leading cybersecurity company protecting organizations' greatest assets and biggest risks: vulnerabilities in people. With an integrated suite of cloud-based solutions, Proofpoint helps companies around the world stop targeted threats, safeguard their data, and make their users more resilient against cyber-attacks. Leading organizations of all sizes, including more than half of the Fortune 1000, rely on Proofpoint for people-centric security and compliance solutions mitigating their most critical risks across email, the cloud, social media, and the web.
We are singularly devoted to helping our customers protect their greatest assets and biggest security risk: their people. That's why we're a leader in next-generation cybersecurity.
Protection Starts with People. Proofpoint.
About the role
We are looking for a skilled Site Reliability Engineer (SRE) to join our team and help build, deploy, scale, and operate highly reliable distributed systems. You will play a key role in deploying and managing platforms across multiple regions, ensuring high availability, scalability, and performance.
Key Responsibilities :
Deploy, manage, and scale distributed platforms across multiple geographic regions
Design and maintain Kubernetes-based infrastructure for large-scale applications Build and manage Helm charts for efficient and repeatable deployments
Monitor system health using Grafana dashboards and metrics; proactively identify and resolve issues Improve system reliability, performance, and scalability through automation and best practices
Handle large-scale deployments and improve infrastructure for growth
Collaborate with development teams to ensure smooth CI/CD and production readiness Implement observability, alerting, and incident response processes Troubleshoot production issues and perform root cause analysis
Write and maintain run books for incident response
Required Qualifications :
4-5 years of experience in Site Reliability Engineering, DevOps, or similar roles Strong hands-on experience with Kubernetes in production environments
Strong experience with infrastructure as code (Terraform, git.)
Strong experience with AWS (eks, vpc, s3, ecr, iam role etc)
Solid experience with Helm charts for application deployment
Strong experience in bash scripting and tooling
Experience with large-scale distributed systems and high-availability architectures Strong understanding of containerization, micro-services, and cloud-native ecosystems
Experience with CI/CD pipelines and automation tools
Good debugging and problem-solving skills in production environments
Preferred Skills:
Proficiency in at least one of the following: Golang, Python.
Experience building and managing Grafana dashboards and metrics
Knowledge of monitoring and observability stacks (Prometheus, Loki, etc.) Experience with multi-region deployments and global infrastructure
Why Proofpoint
Protecting people is at the heart of our award-winning lineup of cybersecurity solutions, and the people who work here are the key to our success. We're a customer-focused and a driven-to-win organization with leading-edge products. We are an inclusive, diverse, multinational company that believes in culture fit, but more importantly 'culture-add', and we strongly encourage people from all walks of life to apply.
We believe in hiring the best and the brightest to help cultivate our culture of collaboration and appreciation. Apply today and explore your future at Proofpoint! #LifeAtProofpoint
**Why Proofpoint?**
At Proofpoint, we believe that an exceptional career experience includes a comprehensive compensation and benefits package. Here are just a few reasons you'll love working with us:
+ Competitive compensation
+ Comprehensive benefits
+ Career success on your terms
+ Flexible work environment
+ Annual wellness and community outreach days
+ Always on recognition for your contributions
+ Global collaboration and networking opportunities
**Our Culture:**
Our culture is rooted in values that inspire belonging, empower purpose and drive success-every day, for everyone.
We encourage applications from individuals of all backgrounds, experiences, and perspectives. If you need accommodation during the application or interview process, please reach out to .
**How to Apply**
Interested? Submit your application along with any supporting information- we can't wait to hear from you!
Proofpoint has been honored with six Best Places to Work Awards in 2024 by workplace culture leader Comparably, including Best Company Career Growth, Best Company Outlook, Best Global Culture, Best Engineering Teams, Best Sales Teams, and Best HR Teams.
We are the leader in human-centric cybersecurity. Half a million customers, including 87 of the Fortune 100, rely on Proofpoint to protect their organizations. We're driven by a mission to stay ahead of bad actors and safeguard the digital world. Join us in our pursuit to defend data and protect people.
Our BRAVE Values:
At Proofpoint, we are BRAVE in everything we do, and our values aren't just words-they shape how we work, collaborate, and grow.
We seek people who are bold enough to challenge the status quo, responsive in the face of ever-evolving threats, and accountable for delivering real impact.
We value those with a visionary mindset who anticipate what's next and push cybersecurity forward, and we celebrate exceptional execution that ensures we continue to defend data and protect people.
Proofpoint is an equal opportunity employer, we hire without consideration to race, religion, creed, color, national origin, age, gender, sexual orientation, marital status, veteran status or disability.
Find your network, your allies, and your biggest fans. We know that work is simply better when you're surrounded by people who inspire you-who share ideas, cheer you on, and genuinely want to see you succeed. That's why we offer social circles, sponsored networks, and connection points across teams and time zones-to help you find your people, build your community, and thrive together.
This isn't just a job-it's a mission to protect people and defend data in a world that never slows down. We're building the future of human-centric cybersecurity, and that future belongs to all of us. We take ownership, move fast, and hold ourselves accountable-because that's what it takes to stay ahead. And we do it together, winning as one.
Be empowered to reach your full potential through meaningful challenges and personalized support-designed around you and your goals. Whether you're growing as a leader or leveling up from great to exceptional as an individual contributor, we're here to help you get there.
Proofpoint is an equal opportunity employer, we hire without consideration to race, religion, creed, color, national origin, age, gender, sexual orientation, marital status, veteran status or disability.
Is this job a match or a miss?
Apply Now

Lead Engineer- Reliability

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

**What you'll do:**
**Key Responsibilities**
+ Work with Product SMEs to influence product design by introducing innovative ideas to mitigate reliability risks and optimize product performance.
+ Serve as an SME and independently deliver Design for Reliability (DfR) activities in electrical product development programs.
+ Develop short-term and long-term technology maturity roadmaps for Design for Reliability of electrical components.
+ Lead Physics of Failure (PoF)-based reliability assessments for electrical/electronic products, as well as cooling system solutions used in data centers.
+ Plan and conduct reliability life testing, leveraging Eaton's internal test laboratories and external agencies as appropriate.
+ Develop data analytics-based solutions using test data from electrical/electronic products to support reliability life modeling.
+ Provide technical mentoring to new talent in handling critical issues and delivering project objectives.
+ Reduce warranty defects through optimized burn-in strategies in close collaboration with manufacturing plants, and drive the adoption of structured problem-solving methodologies such as FRACAS and Shainin Red X for new product development and existing product engineering initiatives.
**A. Technology Responsibilities**
+ Collaborate with business stakeholders to assess current reliability capabilities and identify opportunities for Accelerated Life Testing (ALT) and test optimization.
+ Develop short-term and long-term technology roadmaps to enhance existing reliability technologies and build new organizational capabilities.
+ Drive initiatives to implement risk mitigation strategies across the organization, including the application of structured problem-solving methodologies in new product development and product improvement programs.
**B. Program Responsibilities**
+ Provide technical leadership for the assigned business portfolio through effective stakeholder management and engagement.
+ Plan project milestones, execute deliverables, conduct technical reviews, communicate progress to stakeholders, and support team development through technical mentoring.
+ Create and/or update reliability deliverables, including Reliability Program Plans, Reliability Predictions, Failure Mode and Effects Analyses (FMEA), Highly Accelerated Life Tests (HALT), and Reliability Demonstration Tests (RDT) for new product development programs.
**C. Extended Responsibilities**
+ Lead reliability management (RM)-level initiatives and contribute through impactful value-added activities.
+ Provide technical guidance and mentorship to engineers to ensure high-quality, first-time-right deliverables.
**D. Individual Contributor Responsibilities**
+ Collaborate with cross-functional teams to execute Design for Reliability (DfR) activities for new product development (NPD) programs, primarily focused on data center applications, ensuring the identification, assessment, and mitigation of reliability risks.
+ Understand project requirements and independently solve complex technical challenges through collaboration with local and global stakeholders.
+ Demonstrate innovative thinking and foster a culture of technical rigor to generate intellectual property, including patents, trade secrets, and technical publications.
+ Exhibit a continuous improvement mindset to deliver measurable value in terms of cost, quality, and schedule performance.
**Qualifications:**
**Required:** Bachelor's degree in Electrical Engineering, Electronics Engineering, or Mechanical Engineering, with relevant industry experience.
**Required:** Bachelor's degree with **8+ years** of experience, or Master's degree with **6+ years** of experience, in electrical and/or electronics product reliability engineering.
**Skills:**
**Technical Competencies**
+ Proficiency in Design for Reliability (DfR) processes, including system reliability modeling, reliability goal setting, allocation, prediction, demonstration testing, life data analysis, reliability growth analysis, and Accelerated Life Testing (ALT).
+ Strong knowledge of reliability standards and methodologies, including MIL-HDBK-217F, NPRD, and Telcordia.
+ Experience with FRACAS (Failure Reporting, Analysis, and Corrective Action System) and converting field/warranty data into meaningful life and failure-rate information.
+ **Desired:** Relevant work experience in the reliability engineering of electrical and electronic systems for data center applications.
+ **Preferred:** DFSS Green Belt certification from a recognized organization and ASQ Certified Reliability Engineer (CRE) certification.
+ Experience with industry-standard reliability analysis tools. Preferred tools include Sherlock, ReliaSoft, BlockSim, and Minitab.
**Behavioral Competencies**
+ **Accountability:** Takes full ownership of assigned tasks and consistently delivers results despite challenges and roadblocks.
+ **Passion:** Approaches work with enthusiasm and commitment, inspiring and motivating others to contribute toward common goals.
+ **Resourcefulness:** Collaborates across organizational boundaries and effectively leverages available resources to achieve desired outcomes in a timely manner.
+ **Problem Solving:** Applies structured thinking, sound logic, and proven methodologies to solve complex problems. Looks beyond obvious solutions and develops innovative and practical approaches.
+ **Continuous Learning:** Proactively invests in self-development, continuously enhances skills, and openly shares knowledge and learnings with the broader team.
+ **Transparency:** Fosters an open, collaborative, and trust-based work environment through clear and honest communication.
Is this job a match or a miss?
Apply Now

Sales Engineer

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

Company Description Reliable Terrestrials is a manufacturing company specializing in material handling equipment and industrial safety products. The organization produces safety shoes, safety equipment, rope hoists, hand pallet trucks, and chain hoists for a wide range of industrial applications. Reliable Terrestrials focuses on durable, practical solutions that support safe and efficient operations in warehouses, factories, and construction environments. The company serves customers who require dependable equipment for material movement and workplace safety and aims to build long-term relationships through quality products and support.
Role Description This is a full-time, on-site Sales Engineer role based in Vadodara. The Sales Engineer will work closely with customers to understand their material handling and safety requirements, propose appropriate product solutions, and support them through the sales cycle. Daily tasks include conducting product demonstrations, preparing quotations, responding to technical and commercial inquiries, and coordinating with the production and logistics teams to ensure timely delivery. The role also involves providing basic technical support, gathering feedback from customers, and maintaining strong relationships through regular follow-ups and on-site visits. The Sales Engineer will contribute to achieving sales targets and expanding the company’s presence in the regional market.
Qualifications
  • Candidates should possess strong Sales Engineering and Technical Support skills to understand customer requirements and recommend suitable equipment solutions.
  • Candidates should possess solid Sales and Customer Service skills to manage the sales cycle, address concerns, and build lasting client relationships.
  • Candidates should possess effective Communication skills, including clear verbal and written communication, for interacting with customers and internal teams.
  • Relevant qualifications such as a diploma or degree in engineering (mechanical or related field) or equivalent technical background are beneficial.
  • Experience in industrial equipment, material handling, or safety products, along with familiarity with the Vadodara or surrounding markets, is an advantage.
  • Ability to work on-site, travel to customer locations, use basic office software, and meet defined sales and performance targets.
Is this job a match or a miss?
Apply Now

Dir. Site Reliability Engineering

Noida UKG

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

Why UKG:
At UKG, the work you do matters. The code you ship, the decisions you make, and the care you show a customer all add up to real impact. Today, tens of millions of workers start and end their days with our workforce operating platform. Helping people get paid, grow in their careers, and shape the future of their industries. That's what we do.
We never stop learning. We never stop challenging the norm. We push for better, and we celebrate the wins along the way. Here, you'll get flexibility that's real, benefits you can count on, and a team that succeeds together. Because at UKG, your work matters-and so do you.
UKG is seeking a seasoned Director of Site Reliability Engineering (SRE) to help lead and shape reliability at enterprise scale. You will be responsible for the reliability, resilience, and operational excellence of UKG's platforms worldwide.
This is a high-impact leadership role within a mature, mission-critical environment. You will inherit and lead an established SRE organization responsible for a large, complex and distributed ecosystem comprising hundreds of applications across a hybrid infrastructure spanning public and private clouds.
Success in this role calls for strong systems thinking, operational leadership at scale, and the ability to influence across boundaries. You will drive consistent reliability practices across diverse technologies, modernize how reliability is delivered, and lead globally distributed teams in service of always-on, customer-critical platforms.
**Responsibilities:**
**Production Reliability & Application Behavior**
+ Responsible for reliability outcomes across a large, heterogeneous application portfolio, including availability, performance, scalability, and recoverability
+ Ensure applications meet defined reliability expectations as they operate on both on-prem and cloud platforms
+ Lead and participate in major incident response, acting as a senior escalation point and ensuring effective executive communication
+ Drive post-incident learning and systemic improvements to reduce repeat issues
**Platform-Facing SRE Execution**
+ Lead teams responsible for understanding how applications behave in production, including runtime performance, resource utilization, and failure modes
+ Partner with Infrastructure, Cloud, Security, and Product Engineering teams to address cross-layer reliability concerns
+ Establish standards for operational readiness, release safety, capacity planning, and disaster recovery across platforms
**SRE Practice Consistency at Scale**
+ Apply Site Reliability Engineering principles pragmatically across both legacy and cloud-native systems, including:
+ SLIs, SLOs and reliability targets
+ Error budgets and risk-based decision-making
+ Toil identification and reduction
+ Automation and self-healing where appropriate
+ Observability to support incident response, performance analysis, and proactive capacity management
+ Ensure SRE practices are consistent in intent but adapted in implementation across different technologies and environments
**People Leadership & Organizational Health**
+ Lead and develop SRE managers and engineers across a global organization
+ Inherit existing teams and improve clarity of ownership, execution discipline, and engagement
+ Hire and develop senior SRE leaders capable of operating across both cloud and enterprise platforms
**Strategy, Planning & Influence**
+ Translate business priorities into reliability-focused technical initiatives
+ Partner with senior Product and Engineering leadership to balance delivery velocity, reliability, and operational risk
+ Own and execute against a portion of the SRE roadmap, ensuring transparency, prioritization, and measurable outcomes
+ Advocate for reliability improvements using data, production insight, and operational experience
+ Balance proactive (planned) and reactive work
**Qualifications**
**Required Qualifications**
+ 10+ years of experience in software engineering, systems engineering, SRE, or related disciplines
+ Proven experience leading established, globally distributed engineering organizations
+ Strong understanding of production systems and application behavior at scale
+ Experience operating and leading teams across hybrid environments (on-prem and public cloud)
+ Demonstrated ability to influence outcomes in a matrixed enterprise environment
+ Experience owning incident response, operational reviews, and executive-level communication
+ Excellent communication skills, with the ability to clearly articulate technical and operational concepts to varied audiences
**Preferred Qualifications**
+ Experience supporting large-scale application portfolios across both Windows/.NET and cloud-native environments
+ Familiarity with Google Cloud Platform and enterprise-scale cloud operations
+ Strong understanding of observability practices across application, platform, and infrastructure layers
+ Prior experience partnering closely with Product, Infrastructure, and Cloud leadership
**Company Overview:**
UKG is the Workforce Operating Platform that puts workforce understanding to work. With the world's largest collection of workforce insights, and people-first AI, our ability to reveal unseen ways to build trust, amplify productivity, and empower talent, is unmatched. It's this expertise that equips our customers with the intelligence to solve any challenge in any industry - because great organizations know their workforce is their competitive edge. Learn more at ukg.com.
UKG is proud to be an equal opportunity employer and is committed to promoting diversity and inclusion in the workplace, including the recruitment process.
Disability Accommodation in the Application and Interview Process
For individuals with disabilities that need additional assistance at any point in the application and interview process, please email
It is the policy of Ultimate Software to promote and assure equal employment opportunity for all current and prospective Peeps without regard to race, color, religion, sex, age, disability, marital status, familial status, sexual orientation, pregnancy, genetic information, gender identity, gender expression, national origin, ancestry, citizenship status, veteran status, and any other legally protected status entitled to protection under federal, state, or local anti-discrimination laws. This policy governs all matters related to recruitment, advertising, and initial selection of employment. It shall also apply to all other aspects of employment, including, but not limited to, compensation, promotion, demotion, transfer, lay-offs, terminations, leave of absence, and training opportunities.
Is this job a match or a miss?
Apply Now

IND - Staff Engineer, Reliability

Hyderabad The Hartford

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

IND - Staff Engineer, Reliability - GCC070
We're determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals - and to help others accomplish theirs, too. Join our team as we help shape the future.
**Key Responsibilities**
+ **Data Reliability & Quality:**  Establishand enforce  **Data Service Level Objectives (SLOs)**  focused on data freshness, completeness, and accuracy across critical data products.
+ **Data Observability:**  Implement advanced data observability tools tomonitorthe entire  **data journey** -from ingestion to consumption-detecting data quality anomalies, schema drifts, and pipeline delays in real-time.
+ **Pipeline Resiliency & Automation:**  Collaborate with Data Engineering to embed reliability patterns into data pipelines built using  **Informatica** ,  **Python/** **Pyspark** , and running on platforms like  **Amazon EMR/Hadoop** **, Informatica**  and cloudnativeservices.
+ **Toil Elimination in Data Operations:**  Automate data validation, data reprocessing, data backfilling, and other manual operational tasks within the data lifecycle to reduce toil and improve operational efficiency.
+ **Incident and Problem Management (Data Focus):**  Lead the response and resolution for data-related incidents (e.g., corrupt data, delayed reporting), ensuring fast recovery and effective post-incident reviews (blameless post-mortems).
+ **Runbook Creation & Automation (Data Focus):**  Develop and automate sophisticated, data-aware runbooks for common data pipeline failures, data quality issues, and data recovery scenarios.
**Required Skills & Experience**
+ 8+year'soverall experience in an Infrastructure,Dataor related technology organization with increasing responsibilities as a hands-on technologist.
+ 2-3+ yearexperience in Data Engineering, Data Quality, or a specializedSRErole within an enterprisedataenvironment.
+ Hands-on experience with data warehousing and data lake technologies, including  **Snowflake** , and cloud environments ( **AWS/GCP** ).
+ Hands-on experiencein pipeline development and support using technologies like  **Informatica** ,  **Python/** **Pyspark** , and distributedcompute(EMR/Hadoop).
+ Experience in designing and implementing data quality checks, data validation frameworks, and data governance standards.
+ Hands onexperience in software or cloud engineering. Familiarity with cloud service providersand their core capabilities(compute, containers, databases,APIsetc.).
+ In depth andhands onexperiencewith data observability concepts and tools for monitoring data in motion and at rest (e.g., Monte Carlo,Bigeye,Astro Observe,Datafold, custom solutions).
+ A strong understanding of the "data journey" and the impact of data issues on business outcomes.
+ Expertiseimplementing AIOps tomonitor, manage and self-heal data pipelines, using machine learning principles for anomaly detection.
+ Experience with prompt engineering, implementing AWS or Google AI services,AI enabled automation for data quality,reliabilityand pipeline performance management.
+ Expertisedefining and implementingofDataOpspractices
Is this job a match or a miss?
Apply Now

Site Reliability Manager

Bengaluru Google

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

Site Reliability Manager
_corporate_fare_ Google _place_ Bengaluru, Karnataka, India
**Advanced**
Experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders; deep expertise in domain.
**Minimum qualifications:**
+ Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience.
+ 5 years of experience building or managing distributed systems or cloud infrastructure, with a focus on Kubernetes.
+ 5 years of experience in people management.
+ Experience with site reliability engineering, system design, distributed computing.
**Preferred qualifications:**
+ 5 years of experience in people management, with managing distributed, multi-site teams through engineering managers or tech leads.
+ Experience in Enterprise tooling and technology.
+ Experience in Systems, Applications, and Products (SAP) or other Enterprise Resource Planning (ERP) systems.
**About the job**
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services-both our internally critical and our externally-visible systems-have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE's will keep an ever-watchful eye on our systems capacity and performance.
Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, you'll have the opportunity to manage the complex challenges of scale which are unique to Google Cloud, while using your expertise in coding, algorithms, complexity analysis and large-scale system design. SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.
Core Enterprise System (CES) SRE is part of Corporate Engineering-Site Reliability Engineering (SRE). We provide SRE support to Enterprise applications within Google, powering key verticals such as Finance, Legal, Supply Chain, and HR. Our mission is to deliver service excellence with engineering, innovation and customer focus and transform Google's enterprise domain.
Google is an engineering company at heart. We hire people with a broad set of technical skills who are ready to take on some of technology's greatest challenges and make an impact on users around the world. At Google, engineers not only revolutionize search, they routinely work on scalability and storage solutions, large-scale applications and entirely new platforms for developers around the world. From Google Ads to Chrome, Android to YouTube, social to local, Google engineers are changing the world one technological achievement after another.
**Responsibilities**
+ Manage a team of 6-10 site reliability engineers supporting Google's enterprise services.
+ Develop roadmaps, planning, objectives and key results (OKRs) to move forward the maturity of the managed services. Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation and refinement.
+ Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning and launch reviews. Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
+ Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity.
+ Practice sustainable incident response ensuring services meet their service level objectives.
Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google'sApplicant and Candidate Privacy Policy (./privacy-policy) .
Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See alsoGoogle's EEO Policy ( ,Know your rights: workplace discrimination is illegal ( ,Belonging at Google ( , andHow we hire ( .
If you have a need that requires accommodation, please let us know by completing ourAccommodations for Applicants form ( .
Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.
To all recruitment agencies: Google does not accept agency resumes. Please do not forward resumes to our jobs alias, Google employees, or any other organization location. Google is not responsible for any fees related to unsolicited resumes.
Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.
Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also and If you have a need that requires accommodation, please let us know by completing our Accommodations for Applicants form:
Is this job a match or a miss?
Apply Now

Senior Engineer - Component Reliability

Posted 1 day ago

Job Viewed

Tap Again To Close

Job Description

Achieving our goals starts with supporting yours. Grow your career, access top-tier health and wellness benefits, build lasting connections with your team and our customers, and travel the world using our extensive route network.
Come join us to create what's next. Let's define tomorrow, together.
**Description**
At United, we have some of the best aircraft in the world. Our Technical Operations team is full of aircraft maintenance technicians, engineers, planners, ground equipment and facilities professionals, and supply chain teams that help make sure they're well taken care of and ready to get our customers to their desired destinations. If you're ready to work on our planes, join our Tech Ops experts and help keep our fleet in tip-top shape.
**Job overview and responsibilities**
+ Write concise technical documents for the testing, repair and overhaul of components to comply with manufacturer documents or legal requirements
+ Gather relevant data, report shop findings and initiate corrective action in support of root cause determination on in-service problems or operational issues
+ Ability to manage and decipher large amounts of component data (removal records, shop findings, etc) to determine root cause and possible corrective actions
+ Develops solutions and implementation plans, project justification, cost/benefit analysis and overall management of project implementation which can include obtaining FAA approvals and coordinate warranty recovery on SBs that are applicable
+ As a Component Engineer, you will also evaluate the benefits of the cost impact of a fleet decision to ensure a balance of cost, assets utilization
+ Study, analyze, and seek solutions to production issues related to design, operation, performance, modification, or repair of components
+ Investigate, develop & implement repair processes, procedures for assigned components. Efficiently collaborate with various internal and external resources for developing effective repair schemes for boosting component reliability within acceptable cost parameters. 
+ Capable of interpreting various OEM technical manuals such as, but not limited to: component maintenance manuals (CMM), aircraft maintenance manuals (AMM), fault isolation manuals (FIM), component and airframe service bulletins (SB), etc
+ Provide engineering disposition for technical issues 
+ Work with other engineering departments in evaluation and prototype of new projects
+ Ability to effectively organize and manage multiple priorities for assigned responsibilities
+ Coordinate work with other operational groups to ensure airworthiness, safety, regulatory compliance, operational reliability and operational efficiency
**This position is offered on local terms and conditions. Expatriate assignments and sponsorship for employment visas, even on a time-limited visa status, will not be awarded. This position is for United Airlines Business Services Pvt. Ltd - a wholly owned subsidiary of United Airlines Inc.**
**Qualifications**
**Required Qualifications:**
+ B.E / B.Tech. Degree in Engineering (Aeronautical, Mechanical, Aerospace), related technical field or equivalent relative work experience
+ Minimum of 8 years at an engineer level or similar role elsewhere in the industry
+ Successful candidates will have working knowledge of airline or OEM operations.
+ Engineering skills sets which can be applied to a range of aircraft systems, maintenance programs and engines
+ For service engineering positions, these same technical skills are applied towards operational issues versus project specific work
+ Ability to analyze complex technical issues, detailed level of project management for regulatory compliance modifications
+ Highly detailed project development and management, and overall ownership of specific systems or ATAs.
+ Effective communication skills and strong stakeholder management are essential.
+ Excellent verbal and technical writing skills needed.
+ Must be legally authorized to work in India for any employer without sponsorship
+ Must be fluent in English (written and spoken)
+ Successful completion of interview required to meet job qualification
+ Reliable, punctual attendance is an essential function of the position
**Preferred Qualifications:**
+ Technical Services and reliability Background
+ Work within specific ATA Airline Chapters
+ Airline or Industry experience with general ATA Chapters which could encompass aircraft systems, structures, and disciplines.
+ Proficiency with using database querying tools, such as SQL, Python, MS Office tools (Excel) and data visualization/ reporting tools
+ Possess basic awareness of **Generative AI technologies and tools (such as GPT, Microsoft Copilot, Claude, Google Gemini, and other emerging AI platforms)** to understand their potential applications in improving productivity and technical documentation.
Is this job a match or a miss?
Apply Now