Boost Grafana Projects with Vollna's Tools
Discover how Vollna enhances your Grafana freelancing projects with advanced filters, real-time notifications, and analytics to improve your success on Upwork and other platforms.
Signup for free
to get access to all filter attributes and instant notifications when new jobs are posted.
Setup filter
Get access to over 30+ filter attributes, setup instant notifications, integrate with your CRM and marketing tools, and more.
Start free trial
81 projects
published for past 72 hours.
| Job Title | Budget | Published | |||
|---|---|---|---|---|---|
|
AWS Cloud Migration Architect / Kubernetes Consultant
Applied
|
$35 - $65
/ hr
|
2 hours ago |
Client Rank
- Medium
2 jobs posted
3 open job
Registered: Jul 13, 2026
19:53
3
|
||
|
AWS Cloud Migration Architect / Kubernetes Consultant
Project Overview We are seeking a senior AWS Cloud Migration Architect and Kubernetes Consultant to assess our current dedicated-server infrastructure and lead the design and migration of our platform to AWS. Our current environment includes Kubernetes-based workloads and supporting databases, messaging, API management, identity, security, secrets management, caching, and observability components. The selected consultant will evaluate the current environment, design the target AWS architecture, determine which components should remain self-managed versus move to AWS-managed services, and develop and execute a phased migration plan. This is expected to be an initial 10–12 week consulting engagement, with the possibility of ongoing cloud operations, security, optimization, and architecture support. Engagement Structure The engagement will begin with a paid discovery and architecture phase. The remaining implementation and migration phases will proceed following review and approval of the proposed architecture, migration roadmap, timeline, and cost estimate. The consultant should be prepared to challenge existing assumptions and recommend the best long-term architecture rather than simply migrating the existing environment to AWS without modernization. The estimated timeline below is preliminary and may be adjusted based on the findings of the discovery phase. Phase 1 — Discovery and Assessment Estimated duration: 2 weeks Assess the current dedicated-server and Kubernetes environment. Inventory applications, services, databases, integrations, storage, networking, and dependencies. Review current security, backup, monitoring, deployment, and operational practices. Identify technical risks, migration constraints, and application dependencies. Develop current-state architecture documentation. Recommend the appropriate AWS services and target architecture. Compare AWS-managed services against self-managed deployment options. Prepare a preliminary AWS monthly cost estimate. Produce a phased migration roadmap, risk register, and implementation estimate. Phase 2 — AWS Foundation and Landing Zone Estimated duration: 2–3 weeks Design and implement the AWS account and environment structure. Configure networking, VPCs, subnets, routing, private connectivity, and security boundaries. Establish IAM, role-based access, identity federation, and least-privilege controls. Implement Amazon EKS and supporting Kubernetes services. Establish infrastructure as code using Terraform, CloudFormation, or AWS CDK. Configure container registries, image scanning, secrets management, logging, monitoring, and backup. Establish development, testing, staging, and production environments as appropriate. Implement or improve CI/CD pipelines. Phase 3 — Workload and Data Migration Estimated duration: 4–5 weeks Migrate Rancher or Kubernetes workloads to Amazon EKS. Migrate or modernize databases, messaging, API management, identity, caching, and observability components. Configure storage, load balancing, DNS, networking, encryption, and security controls. Validate application integrations and external dependencies. Perform database and data migration. Conduct performance, security, resilience, and recovery testing. Develop production cutover and rollback plans. Execute the production migration with minimal downtime and data-loss risk. Phase 4 — Stabilization and Knowledge Transfer Estimated duration: 1–2 weeks Monitor and stabilize the production environment. Resolve migration-related performance and reliability issues. Optimize AWS resources and infrastructure costs. Complete architecture diagrams and technical documentation. Deliver deployment, recovery, incident-response, and operational runbooks. Conduct backup restoration and failover testing. Provide knowledge-transfer sessions to the internal technical team. Deliver final recommendations for ongoing cloud operations and improvement. Key Objectives Assess the current infrastructure, applications, dependencies, security controls, and operational requirements. Design a secure, scalable, highly available, and cost-efficient AWS architecture. Migrate existing Rancher and Kubernetes workloads to Amazon EKS. Modernize appropriate components using AWS-managed services. Minimize downtime, operational risk, and data-loss risk. Establish infrastructure automation, monitoring, security, backup, and disaster recovery. Ensure the AWS environment supports SOC 2 controls and future compliance requirements. Establish clear operational ownership, documentation, and support procedures. Responsibilities Conduct a detailed technical assessment of the current infrastructure. Document applications, services, dependencies, databases, messaging, networking, storage, and security requirements. Develop current-state and target-state architecture diagrams. Design the AWS organization, account, environment, and network structure. Design and implement AWS infrastructure, including: Amazon EKS Amazon EC2 Amazon VPC AWS IAM Elastic Load Balancing Amazon Route 53 Amazon S3 Amazon EBS and EFS Amazon ECR Amazon RDS or Aurora AWS Backup Amazon CloudWatch AWS CloudTrail AWS Config AWS Secrets Manager or an approved equivalent Migrate Kubernetes workloads from Rancher or other on-premises Kubernetes environments to Amazon EKS. Evaluate whether Rancher should remain part of the target architecture or be replaced with native AWS and Kubernetes management tools. Build reusable infrastructure using Terraform, CloudFormation, or AWS CDK. Establish separate development, testing, staging, and production environments where appropriate. Implement or improve CI/CD pipelines for applications and infrastructure. Configure container image scanning, vulnerability management, secrets handling, and deployment controls. Design secure connectivity between AWS, internal users, external systems, customer environments, and any remaining infrastructure. Evaluate, migrate, or redesign supporting platforms, including: PostgreSQL MongoDB RabbitMQ Apache Kafka Redis or Garnet OpenSearch WSO2 OpenBao Identity and access-management platforms API gateways and API-management platforms Recommend AWS-managed alternatives when they provide meaningful improvements in reliability, security, scalability, operations, or cost. Implement centralized logging, monitoring, metrics, tracing, and alerting. Define service-level objectives, alert thresholds, escalation procedures, and operational ownership. Design Multi-AZ architectures, backup policies, disaster-recovery procedures, and failover processes. Define recovery-time objectives and recovery-point objectives for critical systems. Incorporate SOC 2-aligned controls, including: Access management Logging and monitoring Change management Encryption Backup and recovery Incident response Evidence retention Vulnerability management Develop detailed AWS cost estimates before implementation. Implement tagging, budgets, alerts, rightsizing, storage lifecycle policies, and ongoing cost controls. Produce migration plans, deployment procedures, rollback plans, runbooks, security documentation, and knowledge-transfer materials. Train and support the internal technical team during and after migration. Required Experience Significant hands-on experience designing and operating production AWS environments. Strong experience with Kubernetes and Amazon EKS. Demonstrated experience migrating production workloads from dedicated servers, private data centers, Rancher, or on-premises Kubernetes environments to AWS. Strong knowledge of AWS networking, including: VPCs Public and private subnets Routing NAT gateways Security groups Network ACLs Private endpoints VPNs Load balancers Strong knowledge of AWS IAM, least-privilege access, role-based access control, service accounts, and identity federation. Experience with infrastructure as code using Terraform, CloudFormation, or AWS CDK. Experience designing CI/CD pipelines for containerized applications. Experience with Docker, container registries, image security, and Kubernetes deployment strategies. Experience operating PostgreSQL and at least one NoSQL database in production. Experience with messaging or event-streaming platforms such as RabbitMQ or Kafka. Experience with centralized logging, monitoring, metrics, tracing, and alerting. Strong understanding of encryption, secrets management, backup, disaster recovery, high availability, and business continuity. Ability to troubleshoot complex application, Kubernetes, network, database, and AWS infrastructure issues. Strong architecture, documentation, communication, and knowledge-transfer skills. Ability to clearly explain recommendations, alternatives, risks, costs, and tradeoffs to technical and executive stakeholders. Preferred Qualifications AWS Solutions Architect Professional, AWS DevOps Engineer Professional, or equivalent AWS certification. Certified Kubernetes Administrator, Certified Kubernetes Application Developer, or equivalent Kubernetes certification. Experience designing AWS environments that support SOC 2 audits. Experience with AWS Organizations, Control Tower, Security Hub, GuardDuty, Inspector, Config, CloudTrail, and centralized security logging. Experience planning and executing low-downtime or zero-downtime production migrations. Familiarity with: OpenSearch WSO2 OpenBao RabbitMQ Apache Kafka PostgreSQL MongoDB Redis Garnet Experience with AWS-managed services such as: Amazon Aurora or Amazon RDS Amazon DocumentDB Amazon MSK Amazon MQ Amazon ElastiCache Amazon OpenSearch Service Amazon Managed Service for Prometheus Amazon Managed Grafana Expected Deliverables Architecture and Planning Current-state infrastructure assessment Application and dependency inventory Security and operational gap assessment Current-state and target-state architecture diagrams AWS account and network design AWS service recommendations Managed-service versus self-managed comparison Preliminary monthly AWS cost estimate Migration risk register Migration roadmap with phases, priorities, dependencies, and timeline Detailed implementation estimate following discovery AWS Platform Implementation AWS account and environment structure VPC and network architecture IAM and access-control framework Amazon EKS implementation Infrastructure-as-code repositories CI/CD framework Logging, monitoring, and alerting Security controls and secrets management Backup and disaster-recovery configuration Development, staging, and production environments Migration and Validation Workload migration Database and data migration Integration validation Performance and load testing Security testing Backup and recovery testing Failover testing Production cutover plan Rollback plan Post-migration validation Documentation and Knowledge Transfer Final architecture diagrams Infrastructure documentation Deployment and rollback procedures Operational runbooks Incident-response procedures Backup and disaster-recovery runbooks Cost-management recommendations Knowledge-transfer sessions Recommendations for ongoing support and operational ownership Estimated Engagement Expected duration: Approximately 10–12 weeks Estimated consulting effort: Approximately 250–350 hours Engagement type: Consulting contract, with the possibility of extension Initial commitment: Paid two-week discovery and architecture phase Continuation: Remaining phases subject to approval of the architecture, migration roadmap, timeline, and implementation estimate The exact duration and level of effort may be adjusted based on the complexity, quality, and documentation of the current environment. Proposal Requirements Please include the following in your proposal: Examples of similar AWS and Amazon EKS migrations you personally led. A description of the scale and complexity of those environments. Your specific role in architecture, implementation, migration, and ongoing operations. Experience migrating Rancher-managed Kubernetes environments. Experience with PostgreSQL, MongoDB, RabbitMQ, Kafka, WSO2, OpenSearch, and secrets-management platforms. Your recommended approach for assessing and migrating our environment. Your approach to minimizing downtime, data-loss risk, and operational disruption. Your approach to SOC 2-aligned AWS architecture and operational controls. Your approach to AWS cost estimation and optimization. An example architecture diagram, migration plan, or runbook with confidential information removed. Relevant AWS and Kubernetes certifications. Your availability during the anticipated 10–12 week engagement. Your hourly rate, weekly availability, and proposed engagement structure. Any assumptions or information required to provide an accurate implementation estimate. Consultant Selection Process Shortlisted consultants may be asked to participate in a technical interview covering: A previous AWS or EKS migration they personally led. Kubernetes architecture and troubleshooting. AWS networking and IAM design. Database and data-migration strategy. Backup, disaster recovery, and failover design. SOC 2-related cloud controls. Cost estimation and optimization. Migration sequencing, validation, cutover, and rollback planning. Client's questions:
Hourly rate:
35 - 65 USD
2 hours ago
|
|||||
|
Senior Java Developer (Kafka & Spring Boot)
Applied
|
$3 - $300
/ hr
|
2 hours ago |
Client Rank
- Excellent
$3 013 total spent
24 hires, 10 active
53 jobs posted
45% hire rate,
4 open job
3.92 /hr avg hourly rate paid
69 hours paid
Registered: Sep 21, 2020
Toronto
20:53
5
|
||
|
About Us
We are a Canadian technology company transforming the marketing industry through AI-driven solutions. Our platform helps businesses automate workflows, process large-scale data, and deliver intelligent customer experiences. We are looking for passionate engineers who enjoy solving complex problems and building scalable systems that power modern AI applications. About the Role We are seeking a Senior Java Developer with strong expertise in Java, Spring Boot, and Apache Kafka to join our engineering team. You will design and develop scalable microservices, build event-driven systems, and contribute to AI-powered products used by enterprise customers. You'll work closely with product managers, AI engineers, and DevOps to deliver reliable, high-performance solutions. Responsibilities Design, develop, and maintain backend services using Java and Spring Boot. Build and optimize event-driven architectures with Apache Kafka. Develop secure and scalable REST APIs and microservices. Integrate with SQL and NoSQL databases. Collaborate with AI engineers to build data pipelines and backend services that support AI-powered features. Write clean, maintainable, and well-tested code following engineering best practices. Monitor, troubleshoot, and optimize application performance in production. Participate in architecture discussions, code reviews, and technical planning. Work closely with Product, QA, DevOps, and Engineering teams in an Agile environment. Required Qualifications 5+ years of professional Java development experience. Strong experience with Java 17+ and Spring Boot. Solid understanding of Spring Framework, REST APIs, and microservice architecture. Hands-on experience with Apache Kafka or other event-streaming platforms. Experience designing distributed systems and asynchronous messaging solutions. Strong knowledge of relational databases (PostgreSQL, MySQL, or Oracle). Experience with Git, Maven or Gradle, and CI/CD pipelines. Familiarity with Docker and Kubernetes. Strong analytical, problem-solving, and communication skills. Nice to Have Experience building AI or machine learning platforms. Experience with cloud platforms such as AWS, Azure, or Google Cloud. Knowledge of Redis, Elasticsearch, or RabbitMQ. Experience with monitoring and observability tools such as Prometheus, Grafana, or Splunk. Experience working in Agile/Scrum teams. Technologies Backend: Java, Spring Boot, Spring Cloud, REST APIs, Microservices Messaging: Apache Kafka, Kafka Streams Databases: PostgreSQL, MySQL, Oracle, Redis Cloud & DevOps: Docker, Kubernetes, Jenkins, Git, Maven, Gradle, AWS, Azure Practices: CI/CD, Distributed Systems, Event-Driven Architecture, Agile/Scrum Why Join Us? Work on cutting-edge AI and marketing technology. Build products used by enterprise customers. Collaborate with a talented and supportive engineering team. Competitive compensation and opportunities for professional growth.
Hourly rate:
3 - 300 USD
2 hours ago
|
|||||
|
AWS Cloud Migration Architect/ Kubernetes Consultant
Applied
|
$60 - $110
/ hr
|
2 hours ago |
Client Rank
- Medium
2 jobs posted
2 open job
Registered: Jul 13, 2026
19:53
3
|
||
|
AWS Cloud Migration Architect / Kubernetes Consultant
Project Overview We are seeking a senior AWS Cloud Migration Architect and Kubernetes Consultant to assess our current dedicated-server infrastructure and lead the design and migration of our platform to AWS. Our current environment includes Kubernetes-based workloads and supporting databases, messaging, API management, identity, security, secrets management, caching, and observability components. The selected consultant will evaluate the current environment, design the target AWS architecture, determine which components should remain self-managed versus move to AWS-managed services, and develop and execute a phased migration plan. This is expected to be an initial 10–12 week consulting engagement, with the possibility of ongoing cloud operations, security, optimization, and architecture support. Engagement Structure The engagement will begin with a paid discovery and architecture phase. The remaining implementation and migration phases will proceed following review and approval of the proposed architecture, migration roadmap, timeline, and cost estimate. The consultant should be prepared to challenge existing assumptions and recommend the best long-term architecture rather than simply migrating the existing environment to AWS without modernization. The estimated timeline below is preliminary and may be adjusted based on the findings of the discovery phase. Phase 1 — Discovery and Assessment Estimated duration: 2 weeks Assess the current dedicated-server and Kubernetes environment. Inventory applications, services, databases, integrations, storage, networking, and dependencies. Review current security, backup, monitoring, deployment, and operational practices. Identify technical risks, migration constraints, and application dependencies. Develop current-state architecture documentation. Recommend the appropriate AWS services and target architecture. Compare AWS-managed services against self-managed deployment options. Prepare a preliminary AWS monthly cost estimate. Produce a phased migration roadmap, risk register, and implementation estimate. Phase 2 — AWS Foundation and Landing Zone Estimated duration: 2–3 weeks Design and implement the AWS account and environment structure. Configure networking, VPCs, subnets, routing, private connectivity, and security boundaries. Establish IAM, role-based access, identity federation, and least-privilege controls. Implement Amazon EKS and supporting Kubernetes services. Establish infrastructure as code using Terraform, CloudFormation, or AWS CDK. Configure container registries, image scanning, secrets management, logging, monitoring, and backup. Establish development, testing, staging, and production environments as appropriate. Implement or improve CI/CD pipelines. Phase 3 — Workload and Data Migration Estimated duration: 4–5 weeks Migrate Rancher or Kubernetes workloads to Amazon EKS. Migrate or modernize databases, messaging, API management, identity, caching, and observability components. Configure storage, load balancing, DNS, networking, encryption, and security controls. Validate application integrations and external dependencies. Perform database and data migration. Conduct performance, security, resilience, and recovery testing. Develop production cutover and rollback plans. Execute the production migration with minimal downtime and data-loss risk. Phase 4 — Stabilization and Knowledge Transfer Estimated duration: 1–2 weeks Monitor and stabilize the production environment. Resolve migration-related performance and reliability issues. Optimize AWS resources and infrastructure costs. Complete architecture diagrams and technical documentation. Deliver deployment, recovery, incident-response, and operational runbooks. Conduct backup restoration and failover testing. Provide knowledge-transfer sessions to the internal technical team. Deliver final recommendations for ongoing cloud operations and improvement. Key Objectives Assess the current infrastructure, applications, dependencies, security controls, and operational requirements. Design a secure, scalable, highly available, and cost-efficient AWS architecture. Migrate existing Rancher and Kubernetes workloads to Amazon EKS. Modernize appropriate components using AWS-managed services. Minimize downtime, operational risk, and data-loss risk. Establish infrastructure automation, monitoring, security, backup, and disaster recovery. Ensure the AWS environment supports SOC 2 controls and future compliance requirements. Establish clear operational ownership, documentation, and support procedures. Responsibilities Conduct a detailed technical assessment of the current infrastructure. Document applications, services, dependencies, databases, messaging, networking, storage, and security requirements. Develop current-state and target-state architecture diagrams. Design the AWS organization, account, environment, and network structure. Design and implement AWS infrastructure, including: Amazon EKS Amazon EC2 Amazon VPC AWS IAM Elastic Load Balancing Amazon Route 53 Amazon S3 Amazon EBS and EFS Amazon ECR Amazon RDS or Aurora AWS Backup Amazon CloudWatch AWS CloudTrail AWS Config AWS Secrets Manager or an approved equivalent Migrate Kubernetes workloads from Rancher or other on-premises Kubernetes environments to Amazon EKS. Evaluate whether Rancher should remain part of the target architecture or be replaced with native AWS and Kubernetes management tools. Build reusable infrastructure using Terraform, CloudFormation, or AWS CDK. Establish separate development, testing, staging, and production environments where appropriate. Implement or improve CI/CD pipelines for applications and infrastructure. Configure container image scanning, vulnerability management, secrets handling, and deployment controls. Design secure connectivity between AWS, internal users, external systems, customer environments, and any remaining infrastructure. Evaluate, migrate, or redesign supporting platforms, including: PostgreSQL MongoDB RabbitMQ Apache Kafka Redis or Garnet OpenSearch WSO2 OpenBao Identity and access-management platforms API gateways and API-management platforms Recommend AWS-managed alternatives when they provide meaningful improvements in reliability, security, scalability, operations, or cost. Implement centralized logging, monitoring, metrics, tracing, and alerting. Define service-level objectives, alert thresholds, escalation procedures, and operational ownership. Design Multi-AZ architectures, backup policies, disaster-recovery procedures, and failover processes. Define recovery-time objectives and recovery-point objectives for critical systems. Incorporate SOC 2-aligned controls, including: Access management Logging and monitoring Change management Encryption Backup and recovery Incident response Evidence retention Vulnerability management Develop detailed AWS cost estimates before implementation. Implement tagging, budgets, alerts, rightsizing, storage lifecycle policies, and ongoing cost controls. Produce migration plans, deployment procedures, rollback plans, runbooks, security documentation, and knowledge-transfer materials. Train and support the internal technical team during and after migration. Required Experience Significant hands-on experience designing and operating production AWS environments. Strong experience with Kubernetes and Amazon EKS. Demonstrated experience migrating production workloads from dedicated servers, private data centers, Rancher, or on-premises Kubernetes environments to AWS. Strong knowledge of AWS networking, including: VPCs Public and private subnets Routing NAT gateways Security groups Network ACLs Private endpoints VPNs Load balancers Strong knowledge of AWS IAM, least-privilege access, role-based access control, service accounts, and identity federation. Experience with infrastructure as code using Terraform, CloudFormation, or AWS CDK. Experience designing CI/CD pipelines for containerized applications. Experience with Docker, container registries, image security, and Kubernetes deployment strategies. Experience operating PostgreSQL and at least one NoSQL database in production. Experience with messaging or event-streaming platforms such as RabbitMQ or Kafka. Experience with centralized logging, monitoring, metrics, tracing, and alerting. Strong understanding of encryption, secrets management, backup, disaster recovery, high availability, and business continuity. Ability to troubleshoot complex application, Kubernetes, network, database, and AWS infrastructure issues. Strong architecture, documentation, communication, and knowledge-transfer skills. Ability to clearly explain recommendations, alternatives, risks, costs, and tradeoffs to technical and executive stakeholders. Preferred Qualifications AWS Solutions Architect Professional, AWS DevOps Engineer Professional, or equivalent AWS certification. Certified Kubernetes Administrator, Certified Kubernetes Application Developer, or equivalent Kubernetes certification. Experience designing AWS environments that support SOC 2 audits. Experience with AWS Organizations, Control Tower, Security Hub, GuardDuty, Inspector, Config, CloudTrail, and centralized security logging. Experience planning and executing low-downtime or zero-downtime production migrations. Familiarity with: OpenSearch WSO2 OpenBao RabbitMQ Apache Kafka PostgreSQL MongoDB Redis Garnet Experience with AWS-managed services such as: Amazon Aurora or Amazon RDS Amazon DocumentDB Amazon MSK Amazon MQ Amazon ElastiCache Amazon OpenSearch Service Amazon Managed Service for Prometheus Amazon Managed Grafana Expected Deliverables Architecture and Planning Current-state infrastructure assessment Application and dependency inventory Security and operational gap assessment Current-state and target-state architecture diagrams AWS account and network design AWS service recommendations Managed-service versus self-managed comparison Preliminary monthly AWS cost estimate Migration risk register Migration roadmap with phases, priorities, dependencies, and timeline Detailed implementation estimate following discovery AWS Platform Implementation AWS account and environment structure VPC and network architecture IAM and access-control framework Amazon EKS implementation Infrastructure-as-code repositories CI/CD framework Logging, monitoring, and alerting Security controls and secrets management Backup and disaster-recovery configuration Development, staging, and production environments Migration and Validation Workload migration Database and data migration Integration validation Performance and load testing Security testing Backup and recovery testing Failover testing Production cutover plan Rollback plan Post-migration validation Documentation and Knowledge Transfer Final architecture diagrams Infrastructure documentation Deployment and rollback procedures Operational runbooks Incident-response procedures Backup and disaster-recovery runbooks Cost-management recommendations Knowledge-transfer sessions Recommendations for ongoing support and operational ownership Estimated Engagement Expected duration: Approximately 10–12 weeks Estimated consulting effort: Approximately 250–350 hours Engagement type: Consulting contract, with the possibility of extension Initial commitment: Paid two-week discovery and architecture phase Continuation: Remaining phases subject to approval of the architecture, migration roadmap, timeline, and implementation estimate The exact duration and level of effort may be adjusted based on the complexity, quality, and documentation of the current environment. Proposal Requirements Please include the following in your proposal: Examples of similar AWS and Amazon EKS migrations you personally led. A description of the scale and complexity of those environments. Your specific role in architecture, implementation, migration, and ongoing operations. Experience migrating Rancher-managed Kubernetes environments. Experience with PostgreSQL, MongoDB, RabbitMQ, Kafka, WSO2, OpenSearch, and secrets-management platforms. Your recommended approach for assessing and migrating our environment. Your approach to minimizing downtime, data-loss risk, and operational disruption. Your approach to SOC 2-aligned AWS architecture and operational controls. Your approach to AWS cost estimation and optimization. An example architecture diagram, migration plan, or runbook with confidential information removed. Relevant AWS and Kubernetes certifications. Your availability during the anticipated 10–12 week engagement. Your hourly rate, weekly availability, and proposed engagement structure. Any assumptions or information required to provide an accurate implementation estimate. Consultant Selection Process Shortlisted consultants may be asked to participate in a technical interview covering: A previous AWS or EKS migration they personally led. Kubernetes architecture and troubleshooting. AWS networking and IAM design. Database and data-migration strategy. Backup, disaster recovery, and failover design. SOC 2-related cloud controls. Cost estimation and optimization. Migration sequencing, validation, cutover, and rollback planning. Client's questions:
Hourly rate:
60 - 110 USD
2 hours ago
|
|||||
|
I want to hire the devops engineer to manage the aws services of the my website
Applied
|
$15 - $35
/ hr
|
6 hours ago |
Client Rank
- Medium
1 jobs posted
100% hire rate,
1 open job
Registered: Jul 30, 2026
Lahore
05:53
3
|
||
|
Need a reliable DevOps engineer to manage, optimize, and secure your AWS infrastructure? I provide end-to-end AWS DevOps services to help businesses deploy applications faster, improve system reliability, reduce operational costs, and maintain high availability.
My Services Include: AWS EC2, VPC, IAM, S3, RDS, Route 53, CloudFront, ELB, and Auto Scaling Docker containerization and Kubernetes (EKS) deployment CI/CD pipeline setup using GitHub Actions, Jenkins, GitLab CI, or AWS CodePipeline Infrastructure as Code (Terraform, CloudFormation, or Ansible) Linux server administration (Ubuntu, Amazon Linux) Nginx and Apache web server configuration SSL certificate installation and domain configuration Monitoring and logging with CloudWatch, Prometheus, and Grafana Backup, disaster recovery, and security hardening Performance optimization and cost optimization Application deployment and ongoing infrastructure maintenance Troubleshooting AWS infrastructure and production issues I focus on building secure, scalable, and highly available cloud environments while following AWS best practices. Whether you need a new cloud infrastructure, migration, automation, or ongoing DevOps support, I can provide a reliable solution tailored to your business needs.
Hourly rate:
15 - 35 USD
6 hours ago
|
|||||
|
Backend Engineer for API Development
Applied
|
not specified | 6 hours ago |
Client Rank
- Medium
2 jobs posted
2 open job
Registered: Mar 30, 2026
19:53
3
|
||
|
Backend Engineer
Role Overview We are looking for an experienced Backend Engineer to accelerate the development of our AI-powered voice conversation platform. This role is focused on building scalable backend services, improving our cloud infrastructure, and delivering production-ready APIs that power real-time AI workflows. You will work closely with the founding team to rapidly ship backend features, improve system reliability, and help build the technical foundation for a highly scalable healthcare product. As one of our early engineers, you will work directly with the founders and play an important role in shaping the product during its earliest stages. This is an opportunity to collaborate closely with a world-class AI and healthcare team, including founders with experience at Google DeepMind, Google, Apple, Meta, Stanford University, and leading healthcare institutions. Your work will have an immediate impact on both the product and our users. Primary Objectives - Design and deliver production-ready backend services and APIs that support our AI-powered clinical platform. - Build and optimize high-performance streaming infrastructure for real-time voice AI and LLM-powered workflows. - Develop secure, scalable authentication and authorization systems, including user identity management and API security. - Design, optimize, and maintain relational and NoSQL databases to ensure high availability, reliability, and performance. - Improve cloud infrastructure, deployment pipelines, and backend architecture to support rapid product iteration and scale. - Integrate third-party services, AI models, and external healthcare systems while maintaining reliability and security. - Monitor, identify, troubleshoot, and resolve backend issues in a timely manner to maintain a stable production environment. - Continuously improve code quality, testing, documentation, and engineering best practices. Responsibilities - Design, develop, and maintain scalable backend services and RESTful APIs. - Build distributed systems that support low-latency, high-throughput AI applications. - Design and optimize database schemas, queries, caching strategies, and data pipelines. - Implement secure authentication and authorization mechanisms (OAuth, JWT, RBAC, etc.). - Collaborate closely with frontend engineers to define APIs and deliver end-to-end product features. - Build integrations with cloud services, third-party APIs, AI infrastructure, and healthcare systems. - Improve application performance, observability, scalability, and reliability through monitoring and optimization. - Participate in code reviews and contribute to technical architecture and engineering best practices. - Write clean, maintainable, well-tested, and well-documented code. Required Qualifications - Bachelor's degree in Computer Science, Engineering, or equivalent practical experience. - 2+ years of professional backend software engineering experience. - Strong proficiency in Python. - Experience building production-grade RESTful APIs using modern backend frameworks such as FastAPI, Flask, or Django. - Strong understanding of distributed systems, concurrency, and scalable backend architecture. - Experience with relational databases (PostgreSQL, MySQL) and NoSQL technologies (MongoDB, Redis, etc.). - Experience designing efficient database schemas, indexing strategies, and performance optimization. - Experience with cloud platforms. - Familiarity with Docker, containerized deployments, and CI/CD workflows. - Experience with Git and collaborative software development. - Strong debugging, problem-solving, and communication skills. Nice to Have - Experience building real-time streaming systems (WebSockets, gRPC, Server-Sent Events, Kafka, Redis Streams, etc.). - Experience with authentication systems (OAuth2, OpenID Connect, JWT, RBAC). - Experience with asynchronous programming and task queues (Celery, RabbitMQ, Kafka, etc.). - Experience with monitoring and observability tools such as Prometheus, Grafana, OpenTelemetry, Better Stack, or Datadog. - Experience working with healthcare systems, HIPAA compliance, or healthcare APIs is a plus. Client's questions:
Budget:
not specified
6 hours ago
|
|||||
|
AI Infrastructure Engineer - MCP, LLM Providers, Guardrails, Auth (Gurgaon/NCR preferred)
Applied
|
$10 - $15
/ hr
|
11 hours ago |
Client Rank
- Medium
$165 total spent
6 hires, 2 active
14 jobs posted
43% hire rate,
1 open job
5.00 /hr avg hourly rate paid
5 hours paid
Registered: May 5, 2018
Gurgaon
06:23
3
|
||
|
Looking for an engineer who can both build and rigorously validate AI infrastructure components. Roughly 60% hands-on technical
work, 40% structured testing and written reporting. You will be writing small services, wiring up cloud and provider integrations, driving real traffic through them, and then proving with evidence what works and what does not. We need someone equally comfortable in a terminal and in a written report. LOCATION Strong preference for candidates based in Gurgaon or the wider Delhi NCR area. The work is remote day to day, but being in the same city makes occasional in-person working sessions possible and keeps coordination simple. Elsewhere in India will be considered for the right person. We are not looking for agencies or teams — this is for an individual freelancer. WHAT THE WORK INVOLVES - Writing small purpose-built services in Node.js or Python: MCP servers, webhook endpoints, mock upstreams, agent harnesses - Configuring LLM provider integrations and exercising them with real traffic: Gemini, OpenRouter, AWS Bedrock, and local/open-weight models via Ollama - Setting up and exercising guardrail and content-safety services: AWS Bedrock Guardrails, Google Model Armor, Azure Content Safety, Presidio for PII/PHI detection - Standing up identity providers and testing auth end to end: JWT, OAuth 2.1 with PKCE, OIDC, JWKS validation, token introspection, expiry and revocation - Working with Model Context Protocol: transports, tool schemas, tool discovery, and tool-level authorisation - Deploying and debugging on AWS EC2 with Docker Compose — reading container logs, rendered configs and REST APIs to find root causes rather than guessing - Writing test cases, capturing evidence, reproducing failures, classifying severity, and writing defect reports precise enough to survive being challenged - Node.js and Python — comfortable writing small services in either - Docker and Docker Compose, Linux, SSH - AWS: EC2, IAM, ideally Bedrock - Auth: JWT, OAuth 2.1 / PKCE, OIDC, JWKS. Hands-on with at least one IdP (Keycloak, Auth0, Okta, Entra) - HTTP at a low level: curl, headers, status codes, SSE and streaming responses, JSON-RPC - Git and GitHub, including issue and PR discipline - Model Context Protocol experience, or the ability to pick it up quickly STRONG PLUS - LLM APIs across more than one provider, and OpenAI-compatible / Anthropic-compatible endpoint shapes - Agent frameworks: PydanticAI, LangChain, LangGraph - Observability: OpenTelemetry / OTLP, Prometheus, Grafana, ClickHouse - Load testing: k6, wrk - Experience where correctness had real consequences — PII/PHI handling, payments, access control HOW WE WORK - Async first. Short written update at the end of each working day, responsive during agreed overlap hours - Hourly, time tracked - Conventions and structure already in place — you will be working inside an existing setup, not starting from scratch - Long-term engagement if it goes well WHAT WE WEIGH AS HEAVILY AS SKILLS - Clear written English - Evidence over assertion. "It worked" is not a result - Honesty about uncertainty — "I have not confirmed this yet" is worth more than a confident guess - Consistent daily communication, every working day TO APPLY Answer these briefly in your proposal: 1. Where are you based? 2. Describe a bug or misconfiguration you found where the consequence was serious rather than cosmetic. How did you find it, and how did you prove it? 3. Have you worked with MCP? If yes, what did you build. If no, how would you stand up an MCP server over HTTP and verify it works? 4. Which of these have you configured yourself, hands on: AWS Bedrock, Bedrock Guardrails, Google Model Armor, Azure Content Safety, Presidio, or an IdP? 5. Your hours of reliable overlap with IST. Proposals that do not answer these will not be reviewed.
Hourly rate:
10 - 15 USD
11 hours ago
|
|||||
|
Senior DevOps / Infrastructure Engineer – AI Services Deployment (vLLM, GPUs & Kubernetes)
Applied
|
$250
|
12 hours ago |
Client Rank
- Excellent
$44 241 total spent
89 hires, 12 active
135 jobs posted
66% hire rate,
1 open job
83.30 /hr avg hourly rate paid
67 hours paid
Registered: Aug 24, 2022
hyderabad
06:23
5
|
||
|
We are seeking an experienced DevOps / AI Infrastructure Specialist to design and execute the deployment, scaling, and infrastructure planning for our AI Platform.
Our architecture includes a mix of heavy GPU-based LLM inference services (vLLM, Qwen), vision/OCR models, and lightweight CPU-based microservices. We need an engineer to build a high-availability, auto-scaling deployment pipeline that maximizes GPU utilization while keeping inference latency low. Key Objectives & Responsibilities ::: Infrastructure Design & Orchestration: Set up deployment and autoscaling strategies across hybrid workloads (GPU & CPU). GPU Workload Optimization: Configure dynamic batching and request scheduling for LLMs running on vLLM. Fault Isolation & Independence: Ensure independent service deployments so failure in one service does not impact others. Cost & Resource Efficiency: Implement dynamic autoscaling (e.g., KEDA, HPA, GPU sharing/slicing) to ensure efficient GPU usage without over-provisioning. SERVICE | GPU REQ. | CPU REQ. | DYNAMIC BATCHING | REQ. SCHEDULING | AUTOSCALING -------------------------------------------------------------------------------------------------- TODOZEE - Gemma(vLLM) | Yes | Helper CPUs | Yes | Yes | Yes TODOZEE - Qwen | Yes | No | Yes | Yes | Yes TODOZEE - Surya OCR | Preferred| Yes | No | No | No Call Summary | Yes | Helper CPUs | Yes | Yes | Yes Translation India | Yes | Yes | Optional | Yes | Yes Translation Global | Yes | Minimal | Optional | Yes | Yes Chat Agent | No | Yes | No | CPU Scaling | Yes Voice Agent | No | Yes | No | CPU Scaling | Yes Required Skills & Experience :: Kubernetes & Autoscaling: Strong experience with K8s, Helm, HPA, and event-driven autoscalers (KEDA, Karpenter, or Cluster Autoscaler). AI/LLM Serving Frameworks: Deep hands-on experience with vLLM, Triton Inference Server, or TGI (specifically configuring dynamic batching and request queues). GPU Infrastructure Management: Experience managing NVIDIA GPUs in Cloud environments (AWS EKS, GCP GKE, Azure AKS, or bare-metal providers like RunPod/Lambda Labs). Traffic Management & Routing: Proficiency in Istio, Envoy, or NGINX for low-latency routing and request scheduling. Infrastructure as Code (IaC): Terraform or CloudFormation for repeatable infrastructure. CI/CD & Observability: GitHub Actions/GitLab CI, Prometheus, Grafana, and tracing tools for AI latency monitoring. Deliverables :: Production-ready Infrastructure-as-Code (Terraform/Kubernetes manifests). Autoscaling rules tuned for GPU workloads (handling cold starts vs. latency limits). CI/CD pipelines for independent service deployments. Monitoring dashboard for GPU memory, latency, throughput, and error rates.
Fixed budget:
250 USD
12 hours ago
|
|||||
|
UI/UX Engineer
Applied
|
$15 - $25
/ hr
|
19 hours ago |
Client Rank
- Risky
1 open job
06:23
1
|
||
|
On Grafana, we have a dashboards, need to create a landing page in which the dashboards are embedded. User should be able to view the dashboards in a navigation panel with customised logos and customised design.
Hourly rate:
15 - 25 USD
19 hours ago
|
|||||
|
Backend Developer for Web3 Application(Node.js / TypeScript)
Applied
|
$70 - $120
/ hr
|
1 day ago |
Client Rank
- Medium
1 jobs posted
2 open job
Registered: Jul 27, 2026
Durham
19:53
3
|
||
|
About Us
EasyBlock Innovations is a blockchain development company building enterprise-grade Web3 solutions for startups and businesses worldwide. Our team develops scalable backend systems that power decentralized applications, digital asset platforms, crypto wallets, DeFi protocols, NFT marketplaces, and enterprise blockchain solutions. We are looking for an experienced Backend Developer to build secure, high-performance APIs and backend services that integrate with blockchain networks and modern cloud infrastructure. Responsibilities Design, develop, and maintain scalable backend services and REST/GraphQL APIs. Build secure authentication and authorization systems (JWT, OAuth 2.0). Develop backend services for blockchain and Web3 applications. Integrate smart contracts, wallets, and blockchain nodes into backend systems. Design efficient database schemas and optimize query performance. Implement caching, queues, and background job processing. Build real-time services using WebSockets and event-driven architectures. Write clean, maintainable, and well-tested code. Collaborate with frontend developers, blockchain engineers, and product managers. Monitor application performance and troubleshoot production issues. Participate in code reviews and architectural discussions. Contribute to CI/CD pipelines and cloud deployments. Required Skills 4+ years of professional backend development experience. Strong proficiency in Node.js and TypeScript. Experience with NestJS or Express.js. Solid understanding of RESTful API and GraphQL development. Experience with PostgreSQL, MySQL, or MongoDB. Familiarity with Redis and message queues (RabbitMQ, Kafka, or BullMQ). Experience integrating Web3 libraries such as Ethers.js or Web3.js. Knowledge of Docker and containerized deployments. Experience with Git, GitHub, and Agile development workflows. Strong debugging, problem-solving, and communication skills. Preferred Qualifications Experience with AWS, Google Cloud, or Microsoft Azure. Experience deploying applications with Kubernetes. Knowledge of blockchain architecture and EVM-compatible networks. Familiarity with Solidity and smart contract interactions. Experience building DeFi, NFT, or tokenization platforms. Knowledge of microservices and event-driven systems. Experience with CI/CD tools such as GitHub Actions or Jenkins. Nice to Have Experience with Python or Go. Familiarity with Elasticsearch or OpenSearch. Knowledge of monitoring tools such as Prometheus and Grafana. Experience with blockchain indexing services such as The Graph. Open-source contributions or personal blockchain projects. Project Details Remote Long-term contract 30–40 hours per week Flexible working hours with overlapping collaboration time We're seeking engineers who enjoy building secure, scalable backend systems and are passionate about modern cloud technologies and Web3 infrastructure. If you're excited about solving complex engineering challenges and delivering production-ready software, we'd love to hear from you. Client's questions:
Hourly rate:
70 - 120 USD
1 day ago
|
|||||
|
Optimización WPO WooCommerce y Seguridad
Applied
|
~854 - 1,708 USD
|
1 day ago |
Client Rank
- not enough data
-
|
||
|
Tengo una tienda WooCommerce en plena producción que ya recibe picos de tráfico altos y necesito llevar su rendimiento y seguridad a nivel top.
Objetivos clave • Tiempo de carga inicial por debajo de 1 s, con especial énfasis en minimizar scripts y servir imágenes optimizadas sin perder calidad. • WPO integral: auditoría, implementación y validación de cachés (Object Cache, Full-Page), compresión, lazy-loading, precarga de recursos críticos y revisión de consultas SQL/PHP. • Estabilidad del servidor ante picos: ajuste de Nginx/Apache, PHP-FPM, Redis u OPcache, y monitorización proactiva vía tools como New Relic o Grafana. • Seguridad y mitigación de ataques: hardening WordPress, control de firewall, reglas WAF, CAPTCHAs, bloqueo de IPs y revisión de logs para detectar patrones de ataque. Lo que espero del entregable 1. Informe inicial con benchmarks Lighthouse/GTmetrix, análisis de bottlenecks y plan de acción. 2. Implementación técnica completa dentro de un entorno de staging y posterior despliegue en producción. 3. Documento final con métricas comparativas antes/después, check-list de seguridad aplicada y guía de mantenimiento. Perfil que valoro – Experiencia senior demostrable en WordPress/WooCommerce, PHP avanzado y administración de sistemas Linux. – Dominio de JavaScript moderno para optimizar y diferir assets sin romper funcionalidades. – Casos de éxito en WPO, alta disponibilidad y hardening. – Trabajo independiente; no agencias de marketing generalistas. Si cumples estos requisitos y puedes aportar referencias, cuéntame tu enfoque y el tiempo estimado para cada fase. Estoy listo para empezar de inmediato. Skills: PHP, JavaScript, WordPress, Apache, Nginx, MySQL, Redis, WooCommerce
Fixed budget:
750 - 1,500 EUR
1 day ago
|
|||||
|
Platform Engineers Needed to Review a Job Simulation Test
Applied
|
$150
|
1 day ago |
Client Rank
- Excellent
$637 932 total spent
1 009 hires, 121 active
1 019 jobs posted
99% hire rate,
6 open job
26.17 /hr avg hourly rate paid
8 011 hours paid
Industry: HR & Business Services
Company size: 100
Registered: Oct 25, 2018
Amsterdam
02:53
5
|
||
|
Hey there!
We're developing a new pre-employment job simulation and need real platform and DevOps engineers to review the materials for technical accuracy and realism, so we can be confident the simulation holds up to expert scrutiny before it goes live. This is a short, remote project, about three hours of work in total. You'll review the simulation across three stages and give us written feedback on whether the incident scenario rings true and whether what we're scoring actually separates strong engineers from weak ones. What you'll do - Confirm the skills and behaviors we target are right for a senior platform or DevOps role, and that the incident scenario is realistic - Review the simulation script, the incident briefing exhibit (system metrics table and log output), the follow up probes and the engineering manager persona for authenticity and difficulty - Review the draft scoring anchors and tell us whether they reflect real performance differences at senior level, adjusting or adding criteria from your own experience What we're looking for - 5+ years of hands-on experience as a Platform Engineer, Site Reliability Engineer or DevOps Engineer - Hands-on Kubernetes and container orchestration in production, not just familiarity - Personal experience responding to incidents involving OOMKill events, container restarts or connection pool exhaustion, ideally while on call - Comfortable with observability tooling such as Prometheus, Grafana or Datadog, and scripting in Python or Bash - A strong eye for detail and a solid command of English - Prior experience reviewing assessment or interview content is a plus, not a must We estimate about 3 hours of focused work across the three stages. If this sounds like you, we'd love to hear from you! Client's questions:
Fixed budget:
150 USD
1 day ago
|
|||||
|
Senior Cloud Infrastructure Architect (Kubernetes + GPU Cloud + HPC)
Applied
|
$7 - $10
/ hr
|
1 day ago |
Client Rank
- Excellent
$12 353 total spent
9 hires, 3 active
36 jobs posted
25% hire rate,
1 open job
9.63 /hr avg hourly rate paid
133 hours paid
Registered: Oct 21, 2025
Abu Dhabi
04:53
5
|
||
|
# AI Cloud Infrastructure Consultant (Kubernetes, GPU Cloud, HPC)
## Project Overview We are building a next-generation AI cloud platform for GPU-based training and inference. We're looking for an experienced cloud infrastructure consultant to help design the architecture of the platform and review key technical decisions. This is **not** a Kubernetes administration role. We are looking for someone who has architected or operated large-scale cloud infrastructure, preferably for AI, GPU, or HPC environments. The engagement will focus on architecture, design reviews, and technical guidance while our engineering team handles implementation. --- ## Project Scope We are seeking guidance on designing and reviewing: * GPU cloud architecture * Kubernetes platform architecture * Control plane and data plane design * GPU scheduling and orchestration * Multi-tenant architecture * High-performance networking (Cilium, eBPF, service mesh) * Storage architecture for AI workloads * Infrastructure security and isolation * Monitoring, logging, and observability * Scalability and reliability best practices --- ## Required Experience We're looking for someone with strong experience in several of the following areas: ### Kubernetes * Production Kubernetes * Kubernetes internals * Operators * Helm * Cluster Autoscaler / Karpenter * CNI plugins ### GPU Infrastructure * NVIDIA GPU infrastructure * CUDA * MIG * GPU scheduling * AI inference platforms * Large-scale model serving ### Networking * Cilium * eBPF * Hubble * Service mesh * InfiniBand or RoCE (preferred) ### HPC * Slurm * HPC clusters * Distributed AI training * Multi-node workloads ### Storage * S3-compatible object storage * Ceph, BeeGFS, Lustre, or similar * Kubernetes Persistent Volumes ### Infrastructure * Linux * Terraform * Docker * KVM virtualization * Infrastructure as Code ### Observability * Prometheus * Grafana * Loki * OpenTelemetry --- ## Nice to Have Experience designing infrastructure similar to: * GPU-as-a-Service platforms * AI cloud providers * LLM inference platforms * Distributed training infrastructure * Multi-tenant cloud platforms Experience with platforms such as: * Nebius * CoreWeave * Lambda Cloud * Crusoe * AWS * Google Cloud * Azure * Oracle Cloud --- ## Deliverables Depending on the engagement, deliverables may include: * Architecture reviews * Technical recommendations * Infrastructure diagrams * Design documentation * Best-practice guidance * Review of implementation plans * Regular architecture consultation sessions --- ## Ideal Consultant You have previously helped build or operate production cloud infrastructure at scale and can advise on architectural decisions rather than day-to-day operations. Experience with GPU infrastructure, AI platforms, or HPC environments is highly preferred. --- ## When Applying Please include: * A brief summary of your relevant experience * Similar cloud infrastructure projects you've worked on * Your experience with Kubernetes, GPU infrastructure, or HPC * Any public architecture blogs, GitHub repositories, talks, or technical publications (if available) * Your availability and hourly rate If you've worked on AI cloud platforms, GPU infrastructure, hyperscale cloud services, or similar projects, we'd love to hear about your experience.
Hourly rate:
7 - 10 USD
1 day ago
|
|||||
|
Senior On-Prem Voice AI Engineer — LiveKit and NVIDIA Blackwell
Applied
|
$500
|
2 days ago |
Client Rank
- Excellent
$25 004 total spent
43 hires, 6 active
53 jobs posted
81% hire rate,
1 open job
10.00 /hr avg hourly rate paid
24 hours paid
Registered: Apr 25, 2013
Mississauga
21:53
5
|
||
|
We need a senior engineer to install, integrate and test a completely on-premises voice AI stack for a production customer-care agent.
Our server has: 2 × NVIDIA RTX PRO 6000 Blackwell GPUs 96 GB VRAM per GPU Self-hosted LiveKit and SIP An existing Python voice-agent runtime No caller audio, transcripts, prompts, customer data or responses may be sent to external inference APIs. Proposed Stack NVIDIA Parakeet for streaming STT Whisper large-v3-turbo as an accuracy fallback Gemma 4 E4B as a fast turn interpreter Gemma 4 26B-A4B as the main LLM Kyutai streaming TTS vLLM or SGLang LiveKit Audio Turn Detector v1-mini Silero VAD and local interruption handling Docker, PostgreSQL, Redis and Qdrant Prometheus, Grafana and OpenTelemetry The final model selection may be adjusted based on measured accuracy and latency. Responsibilities Configure NVIDIA drivers, CUDA, PyTorch and NVIDIA Container Toolkit for Blackwell. Deploy and optimize the local STT, LLM and TTS models. Integrate all services with LiveKit Agents and our Python runtime. Configure self-hosted LiveKit SIP, Redis, TURN, TLS, firewall and media ports. Validate telephone codecs, sample rates and audio resampling. Optimize product-name recognition using local aliases and phonetic matching. Benchmark STT accuracy using real anonymized 8 kHz telephone audio. Measure TTFA, barge-in latency, GPU usage and concurrent-call capacity. Build reproducible Docker containers and production health checks. Configure monitoring, automatic recovery and outbound network restrictions. Provide complete installation and operations documentation. Required Experience Applicants should have hands-on experience with: NVIDIA CUDA and GPU inference NVIDIA Blackwell or recent enterprise GPUs vLLM or SGLang PyTorch and Hugging Face Streaming STT and TTS NVIDIA NeMo or Parakeet Whisper Gemma models LiveKit, WebRTC and SIP Python and FastAPI Docker and Linux Prometheus and Grafana This is not a prompt-engineering or hosted-API project. Deliverables Fully working on-premises voice AI stack Reproducible Docker deployment Pinned dependency and model versions STT accuracy report TTFA and concurrency report Monitoring dashboards Security and outbound-traffic verification Installation, operations and troubleshooting guides Source code committed to our repository Knowledge-transfer session
Fixed budget:
500 USD
2 days ago
|
|||||
|
Full-Stack Developer (Next.js/NestJS/Supabase) — Climate-Tech Marketplace MVP, Phase 0-1
Applied
|
not specified | 2 days ago |
Client Rank
- Medium
$768 total spent
2 hires
3 jobs posted
67% hire rate,
1 open job
20.00 /hr avg hourly rate paid
25 hours paid
Registered: Aug 20, 2025
Hamburg
02:53
3
|
||
|
We are seeking a skilled developer with expertise in Supabase RLS, Next.js, n8n workflow automation, NestJS, and PostgreSQL for platform development. The ideal candidate will have experience in backend technologies and be able to work on complex projects. This is a part-time role with a project scale of medium, expected to last 1 to 3 months.
We are building a B2B marketplace platform connecting companies with vetted energy/climate consultants and implementation partners — think a structured, AI-assisted matching process from initial company analysis through to project completion. The architecture, database schema, API contracts, and all 26 core workflows are already fully specified in a complete internal technical handbook (~250 pages) — your job is to build against that specification, not to design it from scratch. What's already decided (you will not need to make these calls): Full tech stack: Next.js/TypeScript (frontend, Vercel), NestJS/TypeScript (backend), PostgreSQL via Supabase incl. pgvector, n8n (self-hosted workflow orchestration), Cloudflare, Stripe, Mailjet, Docker/Coolify deployment, GitHub Actions CI/CD, Grafana/Prometheus/Loki observability, Hetzner Cloud (EU hosting). Complete database schema definitions, REST API contract (38 endpoints, request/response/roles/error codes), and workflow-by-workflow specification (trigger, steps, error handling, fallback, escalation, logging, monitoring) for all 26 core workflows. Role-based access model (5 roles), Row-Level Security approach, admin action hardening (step-up auth, four-eyes principle). What you will do (phased delivery): Phase 0: Stand up the infrastructure (DB schema, backend skeleton, security layer, CI/CD, monitoring, backup) from the provided spec. Phase 1: Build company/consultant/provider registration and the qualification workflow. Phase 2: Build the AI-assisted analysis, funding-eligibility check, and matching engine integration. Phase 3: Build the offer comparison, award, project execution and completion flow. (Phase 4, learning cycle & GDPR workflows, to be scoped separately after Phase 1-3 are delivered.) Recommended approach: We're open to discussing a leaner initial release — e.g. Phase 0-1 plus a simplified Phase 3 without the AI-matching layer (rule-based/manual matching for launch, AI matching as a fast-follower), and a reduced role model at launch. If you see a faster path to a working marketplace core, tell us in your proposal. Full specification is shared under NDA after an initial screening call — see application questions below. This is a serious, well-documented project: you will receive a complete technical handbook, not a vague brief. We're open to: individual full-stack freelancers or small agencies/teams. Please indicate which you are in your proposal. Budget: Fixed-price milestones per phase, or hourly — open to discuss based on your proposal. Phase 0-3 core process. Duration: Estimated 4 months for a reduced Phase 0-1 launch scope; 6–9 months for the full Phase 0-3 core process. Before contract award: Shortlisted candidates will be asked to complete a small paid test task (2-3 hours, compensated) implementing one endpoint against a sample from our API contract. This lets both sides confirm fit before committing to the full engagement. Please note: Only apply if you meet ALL the criteria listed above under "Ideal candidate." This is an Expert-level engagement — we will not consider applications from beginners, students, or freelancers without proven production experience in this exact stack. Generic or templated proposals will be rejected without response. Ideal candidate: Proven production experience with Next.js, NestJS (or comparable Node.js backend framework), PostgreSQL Experience with Supabase (Auth, Storage, RLS) a strong plus Experience with n8n or comparable workflow-automation tools a strong plus Comfortable working from a detailed written specification and asking precise clarifying questions rather than guessing Clear, proactive English communication (async-friendly, since client is in Europe — time zone overlap is a plus but not mandatory) Skills (tags, max. 10) Next.js, NestJS, TypeScript, PostgreSQL, Supabase, REST API Development, Docker, CI/CD, n8n, Row-Level Security Project length 4 months (reduced Phase 0-1 launch scope); option to extend to 6-9 months for full Phase 0-3 Experience level Expert Screening Questions: To confirm you've read the full job post, please start your proposal with the word "Handbuch." Have you built a production system using Supabase Row-Level Security with multiple user roles? Please share a GitHub link or code sample demonstrating this. This project is built entirely from a pre-written 250-page technical specification (DB schema, API contracts, workflow definitions already fixed). How do you typically work when the architecture is already decided rather than designed by you? What's your process for asking clarifying questions vs. making assumptions? Link your GitHub profile (required, not optional). Which of your public/private repos best demonstrates NestJS + PostgreSQL backend work? Have you used n8n (or Temporal/Zapier/Make) for backend workflow orchestration? If yes, briefly describe one workflow you built. Describe a bug you introduced in production and how you caught/fixed it — what would you change in your testing process next time?
Budget:
not specified
2 days ago
|
|||||
|
Senior AI Backend Engineer Needed
Applied
|
$3,000 - $5,000
|
2 days ago |
Client Rank
- not enough data
-
|
||
|
Our product relies on Large Language Models, and I need a seasoned backend engineer to turn those models into a rock-solid, production-ready service. Your main responsibility is API development: designing, building, and optimising Python-based endpoints that expose LLM features to our web and mobile clients.
You should feel at home writing clean, test-covered code in FastAPI (or a comparable framework) and managing everything that makes an API reliable—auth layers, rate-limiting, logging, CI/CD, and containerisation. Although the core models are LLMs, familiarity with common AI stacks such as TensorFlow, PyTorch or Scikit-Learn will help when we experiment with alternative architectures. Key deliverables • A version-controlled codebase in Python that wraps our existing LLM checkpoints behind REST (or gRPC) endpoints • Dockerfile and deployment scripts for staging and production • Unit and integration tests with ≥90 % coverage, plus concise Swagger/OpenAPI docs • Monitoring hooks (Prometheus/Grafana or similar) so we can track uptime and latency Once the first milestone is live we will iterate together on performance tuning, caching, and scaling strategies, so a proactive approach to profiling and optimisation is essential. If building high-performance AI APIs excites you, let’s talk timelines and dive straight into the repo. Skills: Java, Python, Software Architecture, Git, Docker, API Development, FastAPI, Large Language Model
Fixed budget:
3,000 - 5,000 USD
2 days ago
|
|||||
|
Browser Automation & AI Reply System
Applied
|
$8 - $15
/ hr
|
3 days ago |
Client Rank
- not enough data
-
|
||
|
I need a production-ready system that can run scheduled tasks across many AdsPower browser profiles, each profile locked to its own residential or ISP proxy so sessions stay consistent. The automation itself can be built with Playwright or Puppeteer, whichever you are faster with, but it must launch through the AdsPower API, pick the correct proxy, perform the scripted actions, then hand control back to a Celery/Redis job scheduler for the next run.
On top of that engine, I want an AI microservice that listens for new customer messages, enriches the prompt with live inventory and pricing pulled from Postgres, and sends an appropriate reply through an LLM (OpenAI or Anthropic). Accuracy matters more than creativity here: the reply must reference the customer’s question and the current data in our database. Key points you’ll tackle • AdsPower integration for profile creation, proxy assignment and session handoff • Reliable proxy management (sticky sessions, geo-targeting, rotation logic) • Automation scripts in Playwright/Puppeteer with robust error handling and logging • Celery/Redis scheduling so each profile’s jobs run on a defined cadence without overlap • LLM API calls that generate responses to customer messages and queries, using real-time inventory and pricing pulled via SQL • Postgres schema and queries to expose that business data quickly • Optional but welcome: webhook flow with the GoHighLevel API so messages and replies sync automatically Deliverables 1. Source code for the browser automation module 2. AI reply microservice with environment-configurable keys and endpoints 3. Docker-compose or similar so the whole stack spins up reproducibly 4. README that shows how to add new automated actions, set proxy rules and extend the AI prompts 5. A short recorded demo proving multiple AdsPower profiles run concurrently and the AI returns contextual replies Acceptance criteria • At least 15 concurrent AdsPower sessions run without collision or shared cookies • Average reply latency from message receipt to AI-generated response under 15 sec • Logged success/error metrics visible in Grafana or a simple dashboard • Clean hand-off so future devs can swap AdsPower for GoLogin with minimal code changes If you have already built something similar and can show code or a live example, that will put you at the top of the list. Skills: PHP, Java, Python, Software Architecture, PostgreSQL, API, Automation, Celery
Hourly rate:
8 - 15 USD
3 days ago
|
|||||
|
Principal DevOps / Platform Engineer — Kubernetes, AKS and High Availability
Applied
|
$35 - $90
/ hr
|
3 days ago |
Client Rank
- Excellent
$6 042 854 total spent
193 hires, 83 active
262 jobs posted
74% hire rate,
2 open job
31.92 /hr avg hourly rate paid
177 808 hours paid
Industry: Tech & IT
Company size: 100
Registered: Nov 29, 2015
Dublin
21:53
5
|
||
|
Engagement: Full-time contract, 40 hours per week
Duration: Long-term — this is a build, improve and own role, not a short migration project Location: Remote Timezone: Must overlap at least four working hours with both US Eastern Location preference: Candidates based in the EU/EEA with existing authorization to work or contract there are strongly preferred. We cannot provide visa sponsorship. Start: Immediate About EmpowerID EmpowerID is an established enterprise identity security and Identity Governance and Administration company. We have approximately 120 employees and more than 20 years of experience serving complex enterprise customers. We are building the next generation of the EmpowerID Identity Fabric: a modern microservices platform spanning identity governance, identity provider services, AuthZEN-based authorization, workflow orchestration, identity graph services, analytics and AI-agent governance. Our application stack includes: Python and FastAPI microservices Traefik PostgreSQL and pgvector Redis/Valkey Redpanda/Kafka Neo4j ClickHouse OpenBao S3-compatible object storage Docker Compose Kubernetes, primarily Azure AKS Helm, Terraform/Bicep and GitOps delivery Developers run the complete platform locally through a large Docker Compose environment. We already have Helm charts and Kubernetes/AKS deployment capabilities in varying stages of maturity. This is not a greenfield “convert a Compose file to Kubernetes” assignment. We need a senior hands-on platform engineer to assess what exists, strengthen it, eliminate remaining single points of failure, mature our Helm and infrastructure automation, prove recoverability and high availability, and then take long-term operational ownership of the platform. Our primary SaaS environment runs on Azure AKS. We also need a portable Kubernetes deployment profile for enterprise customers running RKE2-class or comparable on-premises infrastructure. You will work directly with the CEO and the engineering leads. A detailed reference architecture and reviewed operationalization plan already exist. We expect you to execute them carefully, improve them when implementation evidence demands it and challenge assumptions that do not survive contact with production. What you will own 1. Assess and harden the current platform Your first responsibility will be to establish an evidence-based baseline of the current Docker, Helm, Kubernetes and AKS environments. Initial work includes: Review and rationalize existing Helm charts and deployment automation Remove remaining environment-specific and filesystem coupling Rotate credentials and remove hardcoded secrets Eliminate publicly exposed database and administrative ports Correct fail-open readiness and health checks Split cache, session/security-state and coordination workloads where required Pin production images by digest and establish controlled release manifests Validate resource requests, limits, disruption budgets and topology placement Implement and test backup and clean-environment restore procedures Produce an explicit risk register instead of assuming that existing Kubernetes deployment equals high availability 2. Mature the AKS and portable Kubernetes platforms You will refine and complete the platform foundation, including: Terraform or Bicep for AKS, networking, private endpoints, Azure Workload Identity, Key Vault, Front Door and load balancers Reusable Helm library patterns for web services, scalable workers, fenced singleton workers, background processors and migration Jobs Per-service charts composed through Argo CD ApplicationSets or an equivalent GitOps model Ordered database migrations and deployment dependencies without recreating Compose-style depends_on behavior Signed, digest-pinned release manifests with provenance, SBOMs and vulnerability scanning Equivalent portable deployment patterns for RKE2-class and customer-managed Kubernetes environments 3. Improve data-layer availability and recoverability You will help operate and mature the following production data services: PostgreSQL: Azure Database for PostgreSQL Flexible Server and CloudNativePG Redis/Valkey: separate cache and security/coordination planes with appropriate eviction, persistence and HA behavior Redpanda/Kafka: multi-broker operation, replication factors, topic partitioning, ordering, consumer lag and outbox-based recovery OpenBao: Raft HA, external auto-unseal, Kubernetes authentication, policies, Transit and secret lifecycle management Neo4j: clustered or explicitly recoverable graph operation ClickHouse: replicated operation through the Altinity operator and Keeper Object storage: durable application artifacts, backup data and recovery material We do not expect one person to begin as the world’s leading expert in every datastore. We do expect deep production experience with several of them, strong distributed-systems judgment and the ability to become operationally competent with the remainder. 4. Establish secure Kubernetes defaults You will make secure operation the default rather than something each service team must rediscover: Restricted Pod Security Admission Non-root and read-only container patterns NetworkPolicy segmentation Private endpoints for data services Workload Identity instead of static cloud credentials Isolated node pools for untrusted or higher-risk execution Image-signature admission Certificate lifecycle management Least-privilege Kubernetes and Azure RBAC Controlled break-glass procedures with audit evidence 5. Build observable, measurable reliability You will implement and operate: Prometheus and Grafana Loki Tempo OpenTelemetry collectors Service and dependency dashboards SLO and error-budget definitions Multi-window burn-rate alerting Database replication and backup-lag alerts Redpanda consumer-lag alerts Worker lease-age and fencing alerts OpenBao sealed-state alerts Certificate-expiration and key-rotation alerts Capacity forecasting and cost visibility 6. Certify failure behavior High availability will be proven, not declared. You will create and execute repeatable failure drills covering: Pod, node and availability-zone loss PostgreSQL and other state-store failover Ambiguous network partitions Worker termination and lease expiry under load Rolling Kubernetes and application upgrades Schema migration interruption Key and certificate rotation during live traffic Backup restoration into a clean environment Kafka drain, replay and transactional-outbox recovery Application reconnection and retry behavior during failover The results will be measured against ratified RPO, RTO and service-level objectives. Where safety and availability conflict, the behavior and operational decision process must be explicit. 7. Own what you build After the platform is hardened, this becomes an ongoing platform engineering and SRE role. You will own: Production operational readiness On-call participation and incident response Capacity planning Platform and datastore upgrades Backup verification and restore drills Security and certificate rotation Reliability reviews Runbook maintenance Root-cause analysis Automation of recurring operational work Quarterly or agreed failure-drill cadence AI-native engineering is required We expect you to use modern AI coding and reasoning tools fluently as part of your daily engineering workflow. You should be comfortable using tools such as Codex, Claude Code, GitHub Copilot or equivalent systems to: Inspect unfamiliar repositories and infrastructure Draft and refactor Terraform, Helm and automation code Generate test matrices and failure-injection tooling Review manifests and configuration changes Investigate incidents Produce and maintain operational documentation Accelerate repetitive platform work “AI-native” does not mean deploying unreviewed generated infrastructure. You must be able to explain how you: Constrain an AI agent’s access Keep production credentials and customer data out of prompts Review generated changes Validate infrastructure plans before applying them Test generated failure and recovery automation Preserve an auditable human approval boundary Required experience You should have: At least seven years in DevOps, platform engineering, SRE or infrastructure engineering At least five years operating production Kubernetes Hands-on responsibility for a substantial multi-service production platform Experience hardening or migrating a real VM, Compose or early-stage Kubernetes platform into a multi-node or multi-zone production environment Deep Azure and AKS experience, including private clusters, availability zones, Workload Identity, Key Vault and managed PostgreSQL Experience with non-Azure Kubernetes such as RKE2, kubeadm or another customer-managed distribution Strong Terraform or Bicep experience Ability to author reusable Helm library charts, not merely install third-party charts GitOps experience with Argo CD or Flux Experience ordering migrations, Jobs and application rollouts safely Strong Linux, networking, DNS, TLS, storage and Kubernetes troubleshooting skills Production experience operating at least three of the major data services in our stack Hands-on Vault or OpenBao experience, including HA, auto-unseal, Kubernetes authentication and policy design Experience restoring production data into a clean environment Experience with observability, SLOs, alerting and incident response Strong written and spoken English You must also be able to explain: Fencing tokens versus advisory locks Lease expiration under network partitions and process pauses Transactional outboxes Idempotency keys and unknown transaction outcomes At-least-once delivery and idempotent consumers Why “exactly once” should not be used casually Why a scheduled backup is not evidence that recovery works Why Kubernetes replica counts alone do not create high availability Particularly valuable experience The following would be helpful: Python and FastAPI Traefik v3 cert-manager KEDA Prometheus and Grafana rule authoring Cosign, SBOM generation and admission controllers Redpanda or Kafka operator experience CloudNativePG Altinity ClickHouse Operator and Keeper Neo4j clustering Redis or Valkey HA Identity, OAuth, OpenID Connect or authorization systems Enterprise security or regulated-customer environments What success looks like Within the first 30–45 days, we expect: A verified inventory of the current deployment and its failure domains Reproduction of the existing AKS and portable Kubernetes deployments Closure or ownership plans for immediate security and recoverability risks Tested backup and restore evidence for the most critical state stores A prioritized platform backlog tied to measurable operational outcomes Within approximately six months, we expect: Repeatable GitOps deployment into clean environments Hardened and reusable Helm patterns Clear AKS and portable Kubernetes profiles Protected roots of trust Measured state-store recovery behavior Safe singleton-worker execution Observable platform SLOs Scripted failure and restoration drills Runbooks that another qualified engineer can execute Longer term, success means the platform remains secure, recoverable and operable without relying on tribal knowledge. Selection process The process consists of: A technical interview with the CEO and platform leadership. A paid, time-boxed pilot using an isolated copy of the platform. Review of the implementation, evidence, documentation and reasoning—not merely whether the final command succeeded. We value engineers who are precise, verify their own work, communicate clearly and are willing to say: This part of the plan is wrong. Here is the evidence, the risk and the safer implementation. To apply Applications that do not answer the following questions will not be considered: Describe a substantial Compose-, VM- or early-Kubernetes-to-production-Kubernetes project you led. How large was the platform, what failed or surprised you, and what would you do differently? Describe how you would run Vault or OpenBao in HA without distributing static unseal material. How would workloads authenticate, and how would you recover the root of trust? A worker writes to a database and must not execute concurrently with another instance. Describe your lease and fencing design. What happens during a network partition, long process pause, Kubernetes reschedule or database failover? Which of PostgreSQL HA, Redis/Valkey, Kafka/Redpanda, OpenBao/Vault, Neo4j and ClickHouse have you operated in production? Describe the scale and your personal responsibility. Give one concrete example of using an AI coding agent for infrastructure or production operations. What did it produce, how did you verify it and how did you prevent unsafe access or deployment? Where are you located? Do you already have authorization to work or provide contracting services in the EU/EEA? State your timezone, normal working hours, weekly availability, earliest start date and expected hourly or monthly rate.
Hourly rate:
35 - 90 USD
3 days ago
|
|||||
|
Senior DevOps & SRE / Infrastructure Engineer
Applied
|
$100
|
3 days ago |
Client Rank
- Risky
2 jobs posted
2 open job
Registered: Jun 29, 2026
19:53
1
|
||
|
JOB TITLE:
Senior DevOps & SRE / Infrastructure Engineer JOB DESCRIPTION: We are looking for an experienced Senior DevOps & Site Reliability Engineer (SRE) to design, build, secure, and maintain our cloud infrastructure. This role blends DevOps automation with SRE-style reliability practices — we need someone who can both ship infrastructure and keep it running smoothly under real production load. About the Project: We need a hands-on engineer to take ownership of our infrastructure end-to-end: architecture, automation, deployment pipelines, monitoring, and incident response. You'll help us build a system that is scalable, secure, cost-efficient, and resilient. Responsibilities: - Design, deploy, and manage cloud infrastructure (AWS/Azure/GCP) - Build and maintain CI/CD pipelines for automated testing and deployment - Implement Infrastructure as Code (Terraform, CloudFormation, Pulumi, or similar) - Set up and manage container orchestration (Docker, Kubernetes) - Improve system reliability - Set up observability and monitoring (Prometheus, Grafana, Datadog, CloudWatch, ELK, etc.) - Lead incident response, root cause analysis (RCA), and post-mortems - Provide on-call support for production issues as needed - Improve system uptime, performance, and scalability - Strengthen security practices across infrastructure and deployment pipelines - Automate repetitive operational tasks to reduce manual toil - Document infrastructure, runbooks, and incident procedures for the team Requirements: - Senior-level experience in DevOps, SRE, or Cloud Infrastructure Engineering - Strong experience with at least one major cloud provider (AWS, Azure, or GCP) - Proficiency with Infrastructure as Code tools (Terraform, Ansible, or similar) - Experience with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, or similar) - Solid understanding of containerization and orchestration (Docker, Kubernetes) - Experience with monitoring/logging and observability tools (Prometheus, Grafana, Datadog, ELK, etc.) - Scripting/automation skills (Python, Bash, or similar) - Excellent communication skills and ability to work independently - Previous experience working with distributed/remote teams is a plus Nice to Have: - Experience with GitOps tools (ArgoCD, Flux) - Experience with multi-cloud or hybrid cloud environments - Background in chaos engineering or disaster recovery planning Project Type: One-time Estimated Duration: 1 to 2 weeks Experience Level: Senior Budget: Fixed price $100 To Apply: Please include: 1. A brief overview of your relevant DevOps/SRE experience 2. Examples of similar projects you've completed, including any reliability/incident work 3. Your availability for this project 4. Your fixed-price quote for the full scope of work
Fixed budget:
100 USD
3 days ago
|
|||||
|
Data Engineer (Clickhouse / Kubernetes /AWS)
Applied
|
not specified | 3 days ago |
Client Rank
- Medium
6 jobs posted
2 open job
Company size: 2
Registered: Feb 17, 2025
02:53
3
|
||
|
Hi everyone,
for a client of ours, we are looking for a freelance Data Engineer who has experince with Clickhouse (Fokus on Optimization and the Medallion Architecture. Start: ASAP Duration: 3 months with option to extend Workload: 3 days per week (if that works for you) Remote / On-site: Full remote Location: Germany Tasks: Build and operate production-grade ETL pipelines Deploy and operate workloads on Kubernetes (EKS or self-managed), including Helm and Kubernetes manifests Automate infrastructure using Terraform Optimize ClickHouse, including schema design, performance tuning, and query optimization Implement CI/CD pipelines with GitHub Actions Set up monitoring and alerting with Prometheus and Grafana, including definition of SLIs and SLOs Create runbooks for deployment, rollback, and incident handling Ensure structured knowledge transfer to the internal DWH team Requirements & Skills: Strong hands-on experience with AWS, Kubernetes (especially EKS), ClickHouse, and Airflow (Python) Strong Python development skills; TypeScript is used in adjacent services and tools Solid experience with Infrastructure as Code, especially Terraform Strong deployment experience with Helm and/or Kubernetes manifests Proven experience with CI/CD using GitHub Actions Experience in building and operating reliable data platform components in production environments Strong communication skills in English; German is a plus Nice to have: AWS Lambda , Airflow, Moose, Appsmith, or Metabase
Budget:
not specified
3 days ago
|
|||||
|
Comprehensive AWS DevOps Project
Applied
|
~130 - 391 USD
|
3 days ago |
Client Rank
- not enough data
-
|
||
|
1.2048 Project
2. AWS Devops CI CD Project-3-tier 3.ECS Project with CodePipeline 4.RDS 5.VPC 6.modules 7.Mini Project - With actual real time File Structure-terraform 8.Integrate with Slack 9.Integrate with Email using SMTP 10.Monitoring EC2 instances with Promotheus and Grafana 11.MongoExpress and MongoDB Project 12.Blue Green Deployment Project-k8s 13.Cannary Deployment Project-k8s 14.K8S Project-Real time Full-Stack Kubernetes Project 15.https-ingress-project-Setting up HTTPS for application with ACM in K8s 16.User-Based RBAC in Kubernetes - Client Certificate 17.DockerSwarm-jenkins 18.hdfc-website-microservices 19.Dockerfile Task for application Deployment 20.PROJECT - Jenkins, Sonar, Nexus and Ansible Modules deployment to Tomcat 21.MINI PROJECT - How to setup Front end code--ansible 22.Jinja2 Templates in Ansible 23.github-actions 24.Integrate Nexus to Jenkins Pipeline 25.Integrate Tomcat 26.Project on FreeStyle Job - Deployment on Tomcat 27.Maven-project 28.3-tier-mega-project 29.cloudformation-project 30.Serverless Project 31.cloudfront 32.Elasticache and Redshift 33.DMS Project 34.LAMP PROJECT 35.Wordpress website hosting on EC2 36.Storage Gateway,RDS 37.EFS and CLI 38.CORS Task 39.Cloudwatchlogs&Memory 40.Stop & Start EC2 Instances Automatically Using Lambda + EventBridge Scheduler 41.Blue–Green Deployment – AWS (ALB + Target Groups + EC2) Skills: Java, Linux, Amazon Web Services, Node.js, Kubernetes, DevOps, Terraform, CI/CD
Fixed budget:
12,500 - 37,500 INR
3 days ago
|
|||||
|
Lead Azure DevOps Infrastructure Engineer (AZ-104/AZ-204) - Healthcare SaaS
Applied
|
$20 - $25
/ hr
|
3 days ago |
Client Rank
- Excellent
$114 692 total spent
38 hires, 10 active
25 jobs posted
100% hire rate,
1 open job
14.54 /hr avg hourly rate paid
6 899 hours paid
Registered: Aug 13, 2017
Camp Hill
10:53
5
|
||
|
ABOUT US
We're an established healthcare technology company running a multi-portal SaaS platform (web apps, APIs, real-time services, AI integrations) entirely on Microsoft Azure. Our DevOps engineer is moving on, and we're looking for an experienced, certified Azure professional to own our infrastructure, CI/CD and platform reliability. THE ROLE Full-time (40 hrs/week), long-term contract position. You will be our dedicated DevOps engineer and the owner of a mature but actively evolving environment: - CI/CD: Own and improve Azure DevOps pipelines across ~20 repositories (.NET 8 backends, Next.js frontends, containerised services) - Infrastructure: Azure Container Apps, App Services, Azure SQL, Cosmos DB, Key Vault, SignalR Service, Storage, Azure OpenAI / Cognitive Services; some AWS (S3, Textract) - Environments and releases: Manage dev/staging/production promotion, DB migrations, environment configuration, and release coordination with the dev team - Monitoring and reliability: Application Insights, Grafana dashboards, alerting, incident response, cost optimisation - Security: Secret and certificate lifecycle management, service principals, access hygiene, backup/DR - Platform improvement: Drive IaC adoption, deployment automation, and developer-experience improvements over time CERTIFICATION REQUIREMENTS - Required (at least one, must be current): AZ-104 (Azure Administrator Associate) or AZ-204 (Azure Developer Associate) - Strongly preferred: AZ-400 (DevOps Engineer Expert) - Bonus: AZ-305 (Solutions Architect Expert), AZ-500 (Security Engineer Associate), DP-300 (Database Administrator Associate) - Please note: fundamentals-level certifications alone (e.g. AZ-900) are not sufficient for this role EXPERIENCE REQUIREMENTS - 4+ years hands-on Azure experience in production environments - Strong Azure DevOps (Pipelines, Repos, Artifacts) experience, including YAML pipelines and multi-repo setups - Experience deploying and operating containerised .NET applications - Scripting proficiency (PowerShell, Bash or Azure CLI) - Strong written English; comfortable writing runbooks and documentation - Significant overlap with Australian business hours (AEST) PREFERRED - IaC experience (Bicep or Terraform) - Experience in healthcare or other regulated / compliance-sensitive industries - Prior experience as the sole or lead DevOps engineer in a small team WHAT WE OFFER - Stable, full-time long-term engagement: this is a long-term contract role, not a project - Full ownership of the platform with direct access to the development team and leadership - Structured 2-4 week handover with our outgoing engineer - Budget: $20-25/hour USD, 40 hrs/week HOW TO APPLY Please include in your proposal: 1. Your current Microsoft certifications and how they can be verified 2. A short summary of an Azure DevOps pipeline environment you have owned end to end (repo count, stack, what you improved) 3. Your experience with Azure Container Apps and containerised .NET deployments 4. Your timezone and the hours you can reliably overlap with AEST Start your proposal with the word "Handover" so we know you have read this ad. Applications without a current AZ-104 or AZ-204 will not be considered.
Hourly rate:
20 - 25 USD
3 days ago
|
|||||
|
Upload job web into server
Applied
|
$310
|
4 days ago |
Client Rank
- Medium
$233 total spent
4 hires, 1 active
7 jobs posted
57% hire rate,
4 open job
Registered: Mar 31, 2026
MExico
18:53
3
|
||
|
Job Title: Web Deployment Engineer (Website Upload & Server Deployment)
Location: Remote / Hybrid Job Type: Full-Time / Contract Job Summary We are looking for a Web Deployment Engineer responsible for uploading, deploying, and maintaining web applications on production servers. The ideal candidate has experience with Linux servers, web hosting environments, deployment automation, SSL configuration, databases, and troubleshooting production issues to ensure reliable and secure website deployments. Key Responsibilities Upload and deploy web applications to production and staging servers. Configure web servers (Apache, Nginx, or IIS). Set up and manage domains, DNS, and SSL certificates. Configure databases (MySQL, PostgreSQL, SQL Server, etc.). Monitor application performance and server health. Troubleshoot deployment, networking, and server-related issues. Implement backup and disaster recovery procedures. Maintain version control using Git. Automate deployments using CI/CD pipelines where applicable. Coordinate releases with development and QA teams. Ensure website security, updates, and patch management. Document deployment procedures and server configurations. Required Qualifications 2+ years of experience deploying web applications. Strong knowledge of Linux and Windows servers. Experience with Apache, Nginx, or IIS. Proficiency with Git and GitHub/GitLab. Knowledge of FTP, SFTP, SSH, and command-line tools. Experience with MySQL, PostgreSQL, or SQL Server. Understanding of DNS, SSL/TLS, and web hosting. Familiarity with cloud platforms such as Azure, AWS, or Google Cloud. Basic scripting skills (Bash, PowerShell, or Python). Preferred Qualifications Experience with Docker and Kubernetes. Knowledge of CI/CD tools (GitHub Actions, Jenkins, Azure DevOps, GitLab CI). Experience with WordPress, Laravel, Node.js, React, or other web frameworks. Familiarity with monitoring tools such as Grafana, Prometheus, or New Relic. Experience implementing security best practices. Technical Skills Linux Administration Apache / Nginx / IIS Git SSH / FTP / SFTP DNS & SSL Management Docker Kubernetes Azure / AWS / Google Cloud MySQL / PostgreSQL CI/CD Pipelines Bash / PowerShell / Python Soft Skills Strong problem-solving abilities Excellent communication skills Attention to detail Ability to work independently Time management Collaboration across technical teams Nice to Have Experience with Infrastructure as Code (Terraform or Ansible). Knowledge of load balancing and CDN services. Experience managing high-availability production environments. This role is ideal for someone who enjoys ensuring websites and web applications are deployed secu
Fixed budget:
310 USD
4 days ago
|
|||||
|
Kubernetes Cluster Deployment on GPU Nodes
Applied
|
$500
|
4 days ago |
Client Rank
- Risky
1 open job
06:23
1
|
||
|
Kubernetes Cluster Deployment on Bare Metal — with NVIDIA GPUs
Project Overview We are looking for an experienced DevOps / Site Reliability / Kubernetes engineer to design and deploy a production-grade, GPU-enabled Kubernetes cluster on bare-metal hardware. The build and must be delivered end-to-end: from OS installation and low-level hardware configuration through to a fully monitored, storage-backed cluster with automated backups. This is a hands-on infrastructure engagement. Hardware is accessed remotely via IPMI/BMC and SSH — no on-site work required. Scope of Work 1. Bare-Metal Provisioning • Ubuntu 24.04 LTS installation across all nodes. • RAID configuration per our specifications. • Network bonding — dual 100 GbE ports per server bonded via LACP (802.3ad), with matching switch-side port-channel configuration. 2. Kubernetes Cluster • Deploy the latest stable Kubernetes release with a highly available control plane across the 5 management nodes. • Configure container runtime, CNI, and core cluster networking. 3. NVIDIA GPU Enablement • NVIDIA driver installation — for the RTX Pro 6000. • NVIDIA Device Plugin configuration. • NVIDIA GPU Operator configuration. • Validation that GPU workloads schedule and run correctly across all worker nodes. 4. Networking • MetalLB configuration for LoadBalancer-type services. 5. Storage • Ceph cluster providing at least 5 TB usable capacity, exposed to Kubernetes via CSI for dynamic PVC provisioning (Rook-Ceph or standalone Ceph acceptable — state your recommendation). 6. Monitoring & Observability • IPMI / hardware monitoring across all nodes (power, thermals, fans, health). • Cluster monitoring stack (Prometheus + Grafana or equivalent) covering node, GPU, and workload metrics, with dashboards and basic alerting. 7. Backups & Disaster Recovery • Automated, scheduled etcd snapshots. • Kubernetes configuration / resource backups (e.g. Velero or equivalent). • All backups shipped to S3-compatible object storage. Deliverables • A fully operational, validated cluster meeting the scope above. • Documentation and runbooks: architecture overview, provisioning steps, recovery procedures, and day-2 operations. • A handover/knowledge-transfer session with our team. Required Skills & Experience • Proven experience deploying production Kubernetes on bare metal at scale. • Strong Linux systems administration (RAID, network bonding, BMC/IPMI). • NVIDIA GPU on Kubernetes — drivers, device plugin, GPU Operator. • Ceph / Rook and Kubernetes CSI storage. • MetalLB and bare-metal networking. • Prometheus/Grafana observability. • etcd and cluster backup/DR strategies with S3. Nice to Have • Experience with GPU cluster tuning for AI/ML or HPC workloads. • Cluster hardening / security best practices. • Prior work with clusters of comparable node count. To Apply Please include: 1. A brief description of a similar bare-metal / GPU Kubernetes cluster you have built. 2. Your recommended approach and tooling for provisioning, GPU enablement, and storage. 3. Any clarifying questions on the items marked above. 4. Confirmation that you can deliver within the timeline below. Timeline: This is a 1-week engagement — we expect the cluster fully built, validated, and handed over within 7 days of start. Please confirm your availability to work to this schedule and note any dependencies you'd need from us to hit it.
Fixed budget:
500 USD
4 days ago
|
|||||
|
Senior DevOps & SRE / Infrastructure Engineer
Applied
|
$15 - $35
/ hr
|
6 days ago |
Client Rank
- Risky
1 jobs posted
1 open job
Registered: Jun 29, 2026
19:53
1
|
||
|
We are looking for an experienced Senior DevOps & Site Reliability Engineer (SRE) to design, build, secure, and maintain our cloud infrastructure. This role blends DevOps automation with SRE-style reliability practices — we need someone who can both ship infrastructure and keep it running smoothly under real production load.
About the Project: We need a hands-on engineer to take ownership of our infrastructure end-to-end: architecture, automation, deployment pipelines, monitoring, and incident response. You'll help us build a system that is scalable, secure, cost-efficient, and resilient. Responsibilities: - Design, deploy, and manage cloud infrastructure (AWS/Azure/GCP) - Build and maintain CI/CD pipelines for automated testing and deployment - Implement Infrastructure as Code (Terraform, CloudFormation, Pulumi, or similar) - Set up and manage container orchestration (Docker, Kubernetes) - Improve system reliability - Set up observability and monitoring (Prometheus, Grafana, Datadog, CloudWatch, ELK, etc.) - Lead incident response, root cause analysis (RCA), and post-mortems - Provide on-call support for production issues as needed - Improve system uptime, performance, and scalability - Strengthen security practices across infrastructure and deployment pipelines - Automate repetitive operational tasks to reduce manual toil - Document infrastructure, runbooks, and incident procedures for the team Requirements: - Senior-level experience in DevOps, SRE, or Cloud Infrastructure Engineering - Strong experience with at least one major cloud provider (AWS, Azure, or GCP) - Proficiency with Infrastructure as Code tools (Terraform, Ansible, or similar) - Experience with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, or similar) - Solid understanding of containerization and orchestration (Docker, Kubernetes) - Experience with monitoring/logging and observability tools (Prometheus, Grafana, Datadog, ELK, etc.) - Scripting/automation skills (Python, Bash, or similar) - Excellent communication skills and ability to work independently - Previous experience working with distributed/remote teams is a plus Nice to Have: - Experience with GitOps tools (ArgoCD, Flux) - Experience with multi-cloud or hybrid cloud environments - Background in chaos engineering or disaster recovery planning Project Type: One-time Estimated Duration: 1 to 2 weeks Experience Level: Senior Budget: Fixed price $300 To Apply: Please include: 1. A brief overview of your relevant DevOps/SRE experience 2. Examples of similar projects you've completed, including any reliability/incident work 3. Your availability for this project 4. Your fixed-price quote for the full scope of work JOB DESCRIPTION: We are looking for an experienced Senior DevOps & Site Reliability Engineer (SRE) to design, build, secure, and maintain our cloud infrastructure. This role blends DevOps automation with SRE-style reliability practices — we need someone who can both ship infrastructure and keep it running smoothly under real production load. About the Project: We need a hands-on engineer to take ownership of our infrastructure end-to-end: architecture, automation, deployment pipelines, monitoring, and incident response. You'll help us build a system that is scalable, secure, cost-efficient, and resilient. Responsibilities: - Design, deploy, and manage cloud infrastructure (AWS/Azure/GCP) - Build and maintain CI/CD pipelines for automated testing and deployment - Implement Infrastructure as Code (Terraform, CloudFormation, Pulumi, or similar) - Set up and manage container orchestration (Docker, Kubernetes) - Improve system reliability - Set up observability and monitoring (Prometheus, Grafana, Datadog, CloudWatch, ELK, etc.) - Lead incident response, root cause analysis (RCA), and post-mortems - Provide on-call support for production issues as needed - Improve system uptime, performance, and scalability - Strengthen security practices across infrastructure and deployment pipelines - Automate repetitive operational tasks to reduce manual toil - Document infrastructure, runbooks, and incident procedures for the team Requirements: - Senior-level experience in DevOps, SRE, or Cloud Infrastructure Engineering - Strong experience with at least one major cloud provider (AWS, Azure, or GCP) - Proficiency with Infrastructure as Code tools (Terraform, Ansible, or similar) - Experience with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, or similar) - Solid understanding of containerization and orchestration (Docker, Kubernetes) - Experience with monitoring/logging and observability tools (Prometheus, Grafana, Datadog, ELK, etc.) - Scripting/automation skills (Python, Bash, or similar) - Excellent communication skills and ability to work independently - Previous experience working with distributed/remote teams is a plus Nice to Have: - Experience with GitOps tools (ArgoCD, Flux) - Experience with multi-cloud or hybrid cloud environments - Background in chaos engineering or disaster recovery planning Project Type: One-time Estimated Duration: 1 to 2 weeks Experience Level: Senior Budget: Fixed price $300 To Apply: Please include: 1. A brief overview of your relevant DevOps/SRE experience 2. Examples of similar projects you've completed, including any reliability/incident work 3. Your availability for this project 4. Your fixed-price quote for the full scope of work
Hourly rate:
15 - 35 USD
6 days ago
|
|||||
|
AI Platform Engineers (MLOps) - 100% remote, ASAP, 12+ months
Applied
|
not specified | 6 days ago |
Client Rank
- Excellent
$44 671 total spent
7 hires, 7 active
249 jobs posted
3% hire rate,
29 open job
61.76 /hr avg hourly rate paid
583 hours paid
Registered: Jan 5, 2023
Berlin
02:53
5
|
||
|
Dear AI Platform Engineers,
We are currently looking for multiple AI Platform Engineers for one of the leading tech companies in Germany. The position is 100% remote and long-term. QUICK FACTS: - Freelancing/Contracting - Start: ASAP - 100% remote - Duration: long-term 12+ months - Capacity: full-time, part-time also possible - Language: German - German business hours REQUIREMENTS: - Experience in AI platform engineering / MLOps, including building and operating platforms for AI, ML, or GenAI workloads. - Knowledge of AWS, Databricks, Kubernetes, Docker, and Linux. - Experience with Infrastructure as Code and GitOps, using tools such as Terraform, Helm, Argo CD, or Flux. - Familiarity with MLOps / LLMOps concepts such as MLflow, model serving, RAG, prompt management, evaluation, and guardrails. - Understanding of CI/CD, observability, security, governance, scalability, and cost optimization for AI and platform workloads. Are you interested? If so, please send us your proposal including your most recent CV as well as your earliest possible starting date. When sending over your details, please include a self-assessment (0-10, 10 = expert level) in the following areas: - Enterprise AI Platform Engineering - AWS - Databricks - Hybrid cloud and on-premises architectures - Kubernetes - Docker - Linux - Infrastructure as Code and GitOps - Terraform, Helm, Argo CD or Flux - Self-service platforms - Reusable platform components, templates and accelerators - MLOps and LLMOps - MLflow, model serving, RAG, prompt management, evaluation and guardrails - CI/CD for AI, data and platform workloads - AI security and governance - IAM, RBAC, secrets management, encryption, policy as code and audit - Logging - Observability - Prometheus, Grafana, OpenTelemetry or CloudWatch - Scalability, resilience and cost optimisation - Autoscaling, high availability, disaster recovery and FinOps - Software engineering - Programming, particularly Python, as well as languages such as Java, C#, Go or TypeScript - Reference architectures and technical standards for AI and GenAI solutions We are looking forward to hearing back from you. Have a great day! Your wynwood tech Recruiting Team ;-)
Budget:
not specified
6 days ago
|
|||||
|
Senior DevOps Engineer for Voice AI and AI Infrastructure
Applied
|
$20 - $30
/ hr
|
7 days ago |
Client Rank
- Good
$1 025 total spent
1 hires
5 jobs posted
20% hire rate,
4 open job
Company size: 10
Registered: Nov 24, 2024
Dover
06:23
4
|
||
|
We are looking for a highly capable DevOps or Platform Engineer based in Latin America to help us manage, scale, and improve the infrastructure behind our production Voice AI platform.
This is not a traditional CI/CD-focused DevOps role. We need someone with strong experience in cloud infrastructure, Kubernetes, real-time systems, networking, observability, and AI workloads. Our platform handles real-time voice conversations and includes components such as SIP servers, LiveKit, Asterisk, WebSockets, AI model services, GPU infrastructure, Kubernetes, and cloud-native services. Responsibilities Review and improve our existing cloud and Kubernetes architecture Manage production deployments across AWS and Azure Improve reliability, scalability, and fault tolerance Troubleshoot SIP, RTP, WebSocket, networking, and real-time media issues Support Asterisk and LiveKit-based voice infrastructure Set up and improve monitoring, alerting, logging, and incident response Optimize infrastructure for high-concurrency voice workloads Manage GPU-based AI inference workloads Improve autoscaling, resource allocation, and cost efficiency Review security, networking, secrets management, and access controls Investigate production incidents and perform root-cause analysis Automate deployments using Terraform, Helm, GitHub Actions, or similar tools Work closely with our AI, backend, and voice engineering teams Required Experience Strong hands-on experience with Kubernetes, Docker, Helm, and Linux Strong experience with AWS, Azure, or both Experience with Terraform or other infrastructure-as-code tools Experience with Prometheus, Grafana, Loki, OpenTelemetry, or similar tools Strong understanding of networking, load balancing, DNS, TLS, and firewalls Experience managing production systems with high concurrency Strong troubleshooting and root-cause analysis skills Experience with Python, Bash, Go, or another scripting language Good written and spoken English High availability and fast communication during critical incidents Strongly Preferred Experience with Voice AI, conversational AI, or real-time communication platforms Experience with LiveKit, Asterisk, Kamailio, OpenSIPS, FreeSWITCH, or similar systems Understanding of SIP, RTP, WebRTC, STUN, TURN, and WebSockets Experience with GPU workloads, NVIDIA container runtime, and AI inference services Experience scaling LLM, STT, TTS, or machine-learning infrastructure Experience working with distributed systems and low-latency applications Familiarity with Redis, PostgreSQL, Kafka, NATS, or similar technologies Experience with on-premise or private-cloud deployments Engagement Freelance or long-term contract Initially milestone-based Potential for an ongoing engagement Preference for candidates based in Mexico, Colombia, Argentina, Brazil, Chile, or other LATAM countries Must overlap with US working hours Must be responsive and available during agreed production support windows Application Requirements Please include: Your location and time zone Your availability per week Your hourly or monthly rate Examples of production infrastructure you have managed Your experience with Kubernetes and high-concurrency systems Any experience with SIP, WebRTC, LiveKit, Asterisk, Voice AI, or AI infrastructure A brief description of the most difficult production incident you have resolved Whether you are open to completing a paid technical assessment Please do not apply if your experience is primarily limited to basic CI/CD pipelines, website hosting, or standard application deployments. We are looking for someone who can independently investigate complex production infrastructure and real-time communication issues.
Hourly rate:
20 - 30 USD
7 days ago
|
|||||
|
Senior Kubernetes/Platform Engineer (10hrs/week, long-term)
Applied
|
$20 - $45
/ hr
|
7 days ago |
Client Rank
- Excellent
$22 574 total spent
28 hires, 3 active
44 jobs posted
64% hire rate,
1 open job
30.63 /hr avg hourly rate paid
672 hours paid
Registered: Sep 16, 2013
Monheim am Rhein
02:53
5
|
||
|
We are looking for an experienced Cloud Security & Kubernetes Engineer to build and maintain the security foundation of our private cloud platform. You will identify security gaps, implement practical solutions, and continuously improve the security, observability, and resilience of our Kubernetes-based infrastructure.
REQUIRED SKILLS - Strong expertise in cloud-native security for private cloud environments (non-AWS) - Extensive hands-on experience with Kubernetes security and operations - Deep understanding of Kubernetes security best practices, including: - RBAC and access control - Network policies - Secrets management - Admission controllers - Container image security - CI/CD security - Supply chain security - Experience designing and implementing: - Infrastructure monitoring - Security monitoring - Alerting - Threat detection and incident response - Experience securing Linux-based infrastructure and containerized workloads - Familiarity with infrastructure-as-code and GitOps workflows - Good understanding of European data protection requirements (GDPR) and security best practices NICE TO HAVE - Experience with observability platforms such as Prometheus, Grafana, Loki, OpenTelemetry, Falco or similar technologies - Knowledge of service meshes (e.g. Istio or Linkerd) - Experience with MySQL high-availability clusters - Experience securing S3-compatible object storage Frontend development experience (Angular, TypeScript frameworks) to support security-related improvements across the full application stack. OUR ENVIRONMENT - Kubernetes cluster hosting the majority of our workloads - High-availability MySQL database cluster - S3-compatible object storage - CI/CD pipeline We value pragmatic engineers who can balance security with operational simplicity. Rather than delivering a generic security audit, we are looking for someone who can identify realistic improvements, explain trade-offs, and work alongside our team to implement the most valuable recommendations. Ideally, you enjoy working across both infrastructure and application layers and can contribute where necessary, including frontend-related security improvements. Client's questions:
Hourly rate:
20 - 45 USD
7 days ago
|
|||||
|
Senior DevOps Consultant (AWS/Kubernetes/Terraform) – Infrastructure Review & Best Practices
Applied
|
$20
|
7 days ago |
Client Rank
- Risky
1 open job
03:53
1
|
||
|
We’re looking for a Senior DevOps Consultant to review our existing cloud infrastructure and provide recommendations for improving scalability, security, deployment automation, and operational reliability.
Our stack includes AWS, Kubernetes (EKS), Terraform, Docker, GitHub Actions, and monitoring with Prometheus/Grafana. The ideal consultant should have experience designing production-grade infrastructure and be comfortable reviewing CI/CD pipelines, IaC, networking, and security best practices. Deliverables: * Infrastructure architecture review * CI/CD pipeline assessment * Terraform best practices * Kubernetes optimization recommendations * Security and cost optimization report * 1–2 consulting sessions with our engineering team
Fixed budget:
20 USD
7 days ago
|
|||||
|
WildFly L2/L3 Support Engineer
Applied
|
$15 - $22.5
/ hr
|
7 days ago |
Client Rank
- Medium
$439 total spent
6 hires
78 jobs posted
8% hire rate,
5 open job
12.20 /hr avg hourly rate paid
31 hours paid
Registered: Aug 13, 2025
Faridabad
17:53
3
|
||
|
Position Title
L2/L3 Support Engineer - WildFly Application Server About The Role We're seeking experienced L2/L3 Support Engineers with deep hands-on WildFly expertise to provide advanced technical support for enterprise Java applications. You'll diagnose complex issues, escalate strategically, mentor L1 teams, and drive continuous improvement in our support operations. Key Responsibilities Incident Management Investigate and resolve L2/L3 escalations related to WildFly deployments, performance, and stability Reproduce issues in test environments and identify root causes Provide detailed technical analysis and resolution recommendations Manage incident tickets with clear documentation and communication Meet SLA targets for response and resolution times Technical Troubleshooting Analyze WildFly logs, thread dumps, heap dumps, and performance metrics Diagnose deployment failures, classloader issues, memory leaks, and performance bottlenecks Troubleshoot security misconfigurations, SSL/TLS issues, and authentication/authorization problems Identify and resolve cluster communication and failover issues Debug application-server integration problems Knowledge & Documentation Create and maintain detailed troubleshooting guides and knowledge base articles Document recurring issues, workarounds, and permanent solutions Develop runbooks for common support scenarios Maintain up-to-date configuration templates and best practices L1 Support Leadership Mentor and guide L1 support team on WildFly concepts and troubleshooting methodology Review L1 tickets and provide technical guidance Escalate appropriately with clear context and analysis Participate in knowledge transfer sessions and training Proactive Improvements Identify patterns in support tickets to drive preventive measures Recommend configuration optimizations and performance tuning Collaborate with infrastructure and engineering teams on production improvements Monitor and report on support metrics and system health Required Experience WildFly Expertise Minimum 3-5 years hands-on experience with WildFly and/or JBoss EAP Deep knowledge of WildFly architecture, deployment models, and runtime behavior Hands-on experience troubleshooting WAR/EAR/JAR deployments Proficiency with WildFly CLI and management console Experience managing datasources, connection pools, and resource adapters Support & Operational Skills 2+ years in L2/L3 technical support or similar advanced troubleshooting role Experience working with SLA-driven support environments Strong incident management and problem-solving skills Ability to analyze logs, stack traces, and diagnostic data Experience with ticketing systems and documentation workflows Java & Server Technologies Solid Java knowledge (no deep coding required, but must understand application behavior) Understanding of Jakarta EE specifications and common patterns Knowledge of JVM fundamentals, garbage collection, and memory management Familiarity with web application concepts (Servlets, REST, messaging) Debugging & Analysis Tools Proficiency with JVM diagnostic tools: jstack, jmap, jstat, JConsole, VisualVM Experience reading and interpreting thread dumps and heap dumps Knowledge of log analysis and pattern recognition Ability to use monitoring/alerting tools (Prometheus, Grafana, ELK, CloudWatch, etc.) System Administration Comfortable with Linux/Unix command-line operations Experience with application server logs and configuration files Basic understanding of networking, DNS, and HTTP Familiarity with container/Kubernetes environments (Docker preferred) Soft Skills Strong communication—explaining complex issues to technical and non-technical stakeholders Patient, methodical approach to problem-solving Ability to work under pressure and handle multiple concurrent issues Proactive learner staying current with WildFly/Java ecosystem Team player willing to mentor junior support staff Attention to detail and thorough documentation habits Customer-focused mindset with commitment to SLA compliance
Hourly rate:
15 - 22.5 USD
7 days ago
|
|||||
|
SRE Guidance: Monitoring and Performance
Applied
|
~6 - 7 USD
|
7 days ago |
Client Rank
- not enough data
-
|
||
|
I am handling the day-to-day reliability of a growing production stack and need a seasoned Site Reliability Engineer to coach me through two key domains: monitoring/alerting and performance optimization. I already manage the basic upkeep, but I want to elevate my skills so that incidents are caught sooner and services run leaner.
Here is what I’m hoping for: • Regular screen-sharing or video sessions where we review my existing dashboards and alert rules, refine the signal-to-noise ratio, and discuss industry best practices. • Deep-dive walkthroughs on performance tuning—profiling services, interpreting latency metrics, and translating findings into configuration or code changes. • Actionable take-home steps after each session so I can apply what we discuss, then bring results back for feedback. I work primarily in a Linux/containerized environment with common open-source tooling, but I’m open to adopting whatever stack you recommend—Prometheus, Datadog, Grafana, or other fit-for-purpose solutions. To make sure we are a match, please tell me about: • Similar mentoring or advisory roles you’ve done. • Your approach to setting up meaningful alerts without alert fatigue. • A brief example of how you diagnosed and fixed a tricky performance bottleneck. We can start with a short engagement to align on goals; if the collaboration clicks, I’m happy to extend for ongoing guidance. Skills: System Admin, Linux, Network Administration, Internet Security, Alerting, Open Source, Site Reliability Engineering, Containerization
Fixed budget:
600 - 700 INR
7 days ago
|
|||||
|
KNN Programmer
Applied
|
not specified | 8 days ago |
Client Rank
- Risky
3 open job
Industry: Engineering & Architecture
Company size: 2
Registered: Jul 1, 2026
04:53
1
|
||
|
Looking for freelancers of KNX Programmers base on UAE
Budget:
not specified
8 days ago
|
|||||
Related freelance jobs queries: