Find OpenCV Projects on Upwork with Vollna
Boost your freelance business with Vollna. Efficiently find OpenCV projects on Upwork using advanced filters, real-time alerts, and performance analytics.
Signup for free
to get access to all filter attributes and instant notifications when new jobs are posted.
Setup filter
Get access to over 30+ filter attributes, setup instant notifications, integrate with your CRM and marketing tools, and more.
Start free trial
143 projects
published for past 72 hours.
| Job Title | Budget | Published | |||
|---|---|---|---|---|---|
|
Modular PySide6 Object Detection desktop App
Applied
|
~21 - 176 USD
|
2 hours ago |
Client Rank
- not enough data
-
|
||
|
I need a Windows-only Python desktop application that is cleanly split into modules. The stack is fixed: PySide6 for the GUI, OpenCV for image handling, ONNX Runtime with the DirectML EP for inference, and DXCam for high-speed screen capture.
Core behaviour • At start-up the app scans a “Models” folder, lets me pick an ONNX file, then instantiates the matching decoder class so any input / output tensor layout is handled transparently. • Capture, inference, automatic target selection, PID calculations, the main UI, and plug-ins must each live in their own .py file; cross-talk happens only through well-defined interfaces. • Target selection is automatic (based purely on the model’s predictions). Manual controls are not required now, yet keep the UI flexible enough that dropdown lists, selection boxes, and on-screen clicks could be wired in later without restructuring the code. • A simulation mode must exist: when enabled, the computed control values move a dummy target rendered inside a test window that the app spawns itself. No process injection, memory scraping, anti-cheat work-arounds, mouse or controller automation, or any other form of unattended gameplay is permitted. Deliverables (all must be met for acceptance) 1. Fully runnable project folder for Windows 10/11. 2. requirements.txt locked to known-good package versions. 3. step-by-step installation & launch instructions. 4. Well-commented source code, especially the model-adapter interface that links decoders to the core. 5. Example plug-in that does nothing except draw diagnostic overlays. 6. No forbidden automation or bypass methods present anywhere in the codebase. The project is a priority; I’d like to see a first working build as soon as you can manage. Feel free to message me if any detail is unclear before you dive in. Skills: C Programming, Python, Software Architecture, C++ Programming, Artificial Intelligence, Image Processing, OpenCV, Computer Vision, Deep Learning, Desktop Application
Fixed budget:
30 - 250 AUD
2 hours ago
|
|||||
|
Automated MLBB Collage Website
Applied
|
~6 - 8 USD
|
2 hours ago |
Client Rank
- not enough data
-
|
||
|
I want to build a dedicated website for Mobile Legends: Bang Bang players that transforms their skin screenshots into a polished, high-quality collage in just one click. Users simply sign in, upload one or more screenshots of their skin collection, and the system automatically detects, crops, and arranges the skins into a clean, organized grid. Once the collage is generated, users can customize it by adjusting the image order, spacing, background colour, and overall layout before downloading it in high resolution or sharing it directly to social media. The goal is to make creating professional-looking MLBB skin collages fast, effortless
Core workflow • User account & login to save previous collages • Screenshot upload → server-side automated collage generation • Real-time preview with basic customisation tools • One-click download (PNG/JPG) and social-media sharing buttons Tech is flexible, but the solution must be fully responsive and run smoothly on both desktop and mobile. If you choose React, Vue, or another front-end, pair it with a solid back-end (Node, Django, or similar) that can handle image processing libraries such as Pillow or OpenCV. Cloud deployment is preferred so I can push updates quickly. Acceptance criteria 1. Uploads at least 20 screenshots and produces a grid collage in under 10 seconds. 2. Final image resolution remains sharp at 1080 px width or higher. 3. Login, download, customisation, and share functions operate without errors across Chrome, Firefox, Safari, and mobile browsers. Hand-over should include the full source code, a brief setup guide, and deployment instructions so I can maintain the site myself. Skills: JavaScript, Python, Mobile App Development, Django, Node.js, Image Processing, OpenCV, Web Development, Frontend Development, Vue.js
Fixed budget:
600 - 800 INR
2 hours ago
|
|||||
|
Car customization visualizer using classical computer vision
(no generative AI)
Applied
|
$30 - $59
/ hr
|
12 hours ago |
Client Rank
- Risky
1 open job
15:24
1
|
||
|
We build mobile apps for car enthusiasts. We're building a feature that
lets users visualize modifications on a car photo — realistically and deterministically — WITHOUT generative AI (no Stable Diffusion, no FLUX, no diffusion models). We want consistent, repeatable results, which is why we're going the classical CV route instead of generative models. THE MODIFICATIONS WE NEED: 1. Color change - Change the car's body paint to any target color - Must preserve original reflections, highlights, shadows and metallic finish (look like a real photo, not a flat repaint) - Only the body changes; windows, tires, background, grille, lights stay untouched 2. Wheel / rim replacement - Replace the car's wheels with a different rim, given as a separate rim image - Must match the perspective of the wheel in the photo (rim warped correctly onto the elliptical wheel area) - Handle lighting, and ideally the brake caliper / wheel gaps realistically 3. Hood replacement - Replace the hood panel with a different hood design (given as a separate image), warped to the car's perspective, with matching lighting 4. Additive parts (spoilers, body kits) — EXPLORATION - Adding parts that don't exist in the original photo (e.g. a rear spoiler) is harder with pure 2D CV. We're open to your honest assessment: fixed-angle PNG overlay, 3D, or another approach. Tell us what's realistic. This is a discussion, not a hard requirement for the trial. EXPECTED APPROACH (open to your input): - Segmentation / masking - Color manipulation in HSV or LAB space (hue shift, luminance preserved) - Homography / perspective warping for wheels and hood - Poisson blending / seamless compositing - Luminance transfer to keep original lighting on replaced parts Tech: Python, OpenCV. Strong experience in image segmentation, perspective transforms and image compositing is a must. HOW WE WORK: We start with a small PAID TRIAL focused on items 1 and 2 (color + wheel). We'll send you 2-3 real car photos, target colors and a rim image. You deliver the results plus a short note on your approach. If we're happy with the quality, we move to the full project — we have a roadmap of car-editing features and expect ongoing work. PLEASE INCLUDE IN YOUR REPLY: - A relevant example from your past work (image compositing / perspective warping / color editing preferred) - Confirmation you can do this WITHOUT generative AI - Your honest take on which of the 4 items are realistic with classical CV - Your rate and rough timeline for the trial
Hourly rate:
30 - 59 USD
12 hours ago
|
|||||
|
Senior Computer Vision Engineer for Virtual Advertising
Applied
|
not specified | 14 hours ago |
Client Rank
- Good
$5 744 total spent
9 hires, 3 active
15 jobs posted
60% hire rate,
1 open job
Industry: Media & Entertainment
Company size: 2
Registered: Feb 6, 2023
Leigh-on-Sea
13:24
4
|
||
|
Project overview
We are looking for an experienced computer-vision engineer to develop a proof-of-concept system for virtual advertising in recorded football broadcast footage. The initial prototype should replace advertising shown on physical LED perimeter boards with new supplied creative, while keeping the replacement correctly positioned as the broadcast camera moves. This is an early-stage technical feasibility project. The first version will work with recorded footage rather than a live broadcast feed. Initial prototype requirements Using a short sample of football broadcast footage, the system should: Identify and track the visible LED perimeter boards. Replace the existing board content with supplied video or animated graphics. Maintain correct scale, perspective and positioning during camera movement. Mask players, officials and other foreground objects when they pass in front of the advertising. Produce a clean exported 1080p video. Use automated or semi-automated computer vision rather than manual frame-by-frame editing. Be capable of being adapted to additional footage and stadium configurations. The successful freelancer must provide the complete source code, documentation and setup instructions. Required experience Applicants should have strong practical experience in several of the following areas: OpenCV Python and/or C++ PyTorch, TensorFlow or similar frameworks Camera calibration Homography and perspective transformation Semantic or instance segmentation Object detection and tracking Video compositing FFmpeg or GStreamer CUDA or GPU optimisation Real-time or near-real-time video processing Experience with broadcast graphics, augmented reality, virtual production, sports analytics, Unreal Engine, NDI or SDI workflows would be useful but is not essential. Initial deliverable The first paid milestone will involve approximately 30–60 seconds of supplied football footage. The expected deliverable is: Replacement of the visible LED advertising. Stable tracking during camera pan, tilt and zoom. Foreground masking when players cross the advertising area. An exported demonstration video. Complete source code. Instructions explaining how to run the system with replacement creative. This is not a video-editing or manual rotoscoping project. We are looking for a technically reusable computer-vision solution. Potential future work If the prototype is successful, further work may include: Longer football broadcasts. Automated stadium and board calibration. Virtual 3D advertising carpets. Support for different stadiums and camera positions. An operator interface. GPU optimisation. Near-live or live video processing. Multiple regional advertising outputs. There is potential for a longer-term technical relationship for the right person or team. Application questions Please answer the following questions in your proposal: Describe a relevant project where you anchored graphics or virtual objects to footage from a moving camera. How would you track the geometry of LED perimeter boards during camera movement? How would you mask players passing in front of the replacement advertising? Would you use classical computer vision, machine learning or a combination of both? How would you minimise manual calibration or rotoscoping? What processing speed would you expect to achieve at 1080p? What additional development would be required to make the system operate live? Which parts of your previous work can you demonstrate on a video call? Are you personally completing the work, or would any part be subcontracted? Generic AI proposals will not be considered. Please explain your proposed technical approach and include examples of directly relevant work. Commercial terms The project will be structured through fixed-price milestones. The client must receive: Ownership of all bespoke source code following payment. Access to the code throughout development through a client-controlled repository. A list of all third-party libraries, models and licences. Full disclosure of any existing proprietary components. Documentation sufficient for another engineer to run and review the system. The supplied footage and commercial information must remain confidential and may not be used in portfolios or demonstrations without written permission. Please include: Your proposed budget for the initial prototype. Estimated delivery schedule. Recommended technology stack. Any assumptions or technical limitations.
Budget:
not specified
14 hours ago
|
|||||
|
Automated ProductSlide Generator, Template Tool Development
Applied
|
$200
|
16 hours ago |
Client Rank
- Risky
$70 total spent
2 hires
4 jobs posted
50% hire rate,
1 open job
Registered: Oct 11, 2024
Weston
08:24
1
|
||
|
We have an existing tool that places branded product images into a slide template, and we need a developer to refine it, adjust layout, spacing, and watermark placement so the final result looks clean and polished. Python and image processing experience preferred.
Fixed budget:
200 USD
16 hours ago
|
|||||
|
Senior Full-Stack AI Engineer (Python)
Applied
|
$25 - $47
/ hr
|
19 hours ago |
Client Rank
- Risky
1 open job
Registered: Jun 20, 2026
13:24
1
|
||
|
We are a stealth-mode startup building an AI-powered automation platform for the e-commerce and media technology industry.
We're looking for a Senior Full-Stack AI Engineer to develop a standalone AI media processing engine as part of our MVP. This project has a well-defined scope, detailed technical specifications, and the potential for ongoing collaboration after successful delivery. Scope of Work You'll build a standalone Python-based service that: Accepts short user-uploaded videos Extracts and processes video frames using FFmpeg/OpenCV Transcribes audio using Whisper Analyzes visual and audio content with multimodal LLMs (OpenAI GPT-4o or Gemini) Generates structured JSON outputs via REST APIs Delivers clean, production-ready, well-documented code A complete PRD, API specifications, JSON schemas, and architecture documentation will be provided to shortlisted candidates.
Hourly rate:
25 - 47 USD
19 hours ago
|
|||||
|
Senior AI Video Analytics Engineer (Python | FFmpeg | Kafka | NVIDIA Triton | Computer Vision)
Applied
|
$30 - $50
/ hr
|
22 hours ago |
Client Rank
- Risky
1 open job
Registered: Jul 20, 2022
Pune
17:54
1
|
||
|
Job Overview
We are building an enterprise-grade AI Video Intelligence Platform (VMS) that processes hundreds of live CCTV camera streams in real time. We are looking for an experienced Senior Backend / AI Video Analytics Engineer with strong expertise in video streaming, distributed systems, computer vision, and scalable backend architecture. If you have hands-on experience building production-grade video analytics platforms, we would love to work with you. Responsibilities: *Video Streaming & Processing 1. Design and maintain scalable RTSP video ingestion pipelines using FFmpeg and GStreamer 2. Optimize video streams through FPS throttling, frame sampling, motion filtering, and bandwidth optimization 3. Publish video frames and metadata to Apache Kafka for downstream AI processing 4. Ensure high-performance, low-latency processing across hundreds of simultaneous camera streams AI Analytics Pipeline - Develop Kafka consumers for real-time AI inference - Integrate NVIDIA Triton Inference Server - Deploy and optimize ONNX deep learning models - Build face recognition pipelines using: *SCRFD *ArcFace *Qdrant Vector Database - Optimise inference performance for GPU environments Live Streaming & Video Storage 1. Implement HLS live streaming and video restreaming 2. Build continuous recording services using MinIO / Amazon S3-compatible storage 3. Develop: - Video retention policies - Event clip generation - Snapshot services - Video evidence retrieval Platform & Backend Services 1. Develop multi-tenant configuration management 2. Build WebSocket-based real-time alerting services 3. Implement: - Detection visualization - Event exploration - Face gallery management - Event history APIs Infrastructure & DevOps 1. Manage Docker-based deployments 2. Maintain and optimize: - Apache Kafka - PostgreSQL - MinIO - NVIDIA Triton 3. Monitor application performance, health, and scalability 4. Troubleshoot production issues across streaming, AI inference, and distributed services Required Skills - Python (Advanced) - FastAPI / Flask - FFmpeg - GStreamer - RTSP / RTP / HLS - Apache Kafka - Docker - PostgreSQL - Redis - MinIO / Amazon S3 - NVIDIA Triton Inference Server - ONNX Runtime - OpenCV - Computer Vision - Face Recognition (SCRFD, ArcFace) - Qdrant Vector Database - WebSockets - Linux - Git Nice to Have - CUDA optimization - TensorRT - DeepStream SDK - Kubernetes - YOLO models - ByteTrack / DeepSORT - Multi-camera AI systems - Video Management Systems (VMS) - Edge AI deployments Ideal Candidate - 5+ years of backend development experience - 3+ years working with AI video analytics or computer vision - Experience building large-scale distributed systems - Strong understanding of event-driven architectures - Comfortable processing hundreds of concurrent video streams - Able to write clean, production-ready, scalable code Project This is a long-term engagement to build an enterprise AI-powered Video Management System (VMS) for commercial and industrial customers. The right candidate will have the opportunity to become a core technical contributor as the platform scales. To Apply Please include: 1. Links to similar AI Video Analytics or VMS projects. 2. Experience with FFmpeg, GStreamer, Kafka, and NVIDIA Triton. 3. Your experience with face recognition pipelines (SCRFD, ArcFace, etc.). 4. Your preferred tech stack and architecture for processing 500+ RTSP camera streams. 5. Your GitHub or portfolio (if available).
Hourly rate:
30 - 50 USD
22 hours ago
|
|||||
|
Computer Vision Engineer — Player Tracking & Sports Video Analytics (YOLO, ByteTrack)
Applied
|
$30 - $50
/ hr
|
22 hours ago |
Client Rank
- Risky
2 open job
09:24
1
|
||
|
We're building a computer-vision pipeline that turns raw soccer match video
into player tracking data, physical/tactical performance metrics, and opponent scouting profiles, inside a Django-based sports SaaS platform. We need a computer vision engineer to take this from an early prototype to a validated, production-quality pipeline, delivered in sequential phases. CURRENT STATE Player detection runs on YOLO but without proper filtering/validation; tracking is a rough prototype with no occlusion handling; team assignment, real-world coordinate mapping, physical metrics, ball detection, and event detection are either missing or not reliable. SCOPE — 6 PHASES (each independently deployable behind a config flag) - Phase 0: fix detection filtering, real FPS handling, and build an evaluation harness (annotated reference clips + automated scoring) that every later phase is measured against. - Phase 1: integrate a robust multi-object tracker (e.g. ByteTrack) and assign teams via jersey-color clustering with temporal voting. - Phase 2: camera-to-pitch coordinate calibration (homography) so positions can be measured in real meters. - Phase 3: trajectory smoothing and real distance/speed/sprint/acceleration metrics with plausibility checks; remove fabricated team metrics, add the ones genuinely computable (compactness, width, etc.). - Phase 4: ball detection (small, hard object), possession, passes, duels, shots. Highest-uncertainty phase — we plan to start with a time-boxed proof of concept before committing to full scope. - Phase 5: population-based normalization of player/team profiles. - Phase 6: production hardening — adaptive performance tuning, GPU evaluation, QA artifacts visible to end users. You don't need to commit to all 6 phases up front — happy to start with Phase 0–1 (or 0–2) as a paid trial engagement and continue from there. TECH ENVIRONMENT Python, Django, OpenCV, Ultralytics YOLO, Supervision (ByteTrack), SciPy, scikit-learn. Heavy CV code runs inside a Celery worker, isolated from the Django serverless deployment. WHAT WE'RE LOOKING FOR - Proven experience with object detection (YOLO or similar) and multi-object tracking in video. - Experience with camera calibration / homography for pixel-to-real-world coordinate mapping. - Comfortable defining measurable, defensible metrics — validated against ground truth, not "looks about right." - Bonus: sports analytics, small-object detection (ball/puck tracking), or SAHI/tiling techniques. - Python proficiency; able to work inside an existing Django/Celery codebase without introducing heavy imports into the serverless layer. A short technical brief is available attached to this post.
Hourly rate:
30 - 50 USD
22 hours ago
|
|||||
|
AI SMART MOOD DETECTOR
Applied
|
~131 - 392 USD
|
1 day ago |
Client Rank
- not enough data
-
|
||
|
I need a compact Flask web application that taps into the user’s webcam, feeds each frame through OpenCV and TensorFlow, and returns the dominant facial emotion in plain text on screen. I already have the emotion categories pinned down—Happy, Sad, Angry, Surprise, Fear, Disgust and Neutral—so the model should simply detect and display one of those labels without offering any follow-up recommendations.
The interface must be equally comfortable on desktop browsers and mobile devices, which means responsive HTML/CSS and careful handling of camera permissions on smaller screens. Because FER works well for this task, I’d like you to rely on that library alongside the usual OpenCV and TensorFlow stack. The end result should run smoothly on Render, so include a Procfile, requirements.txt and any build commands needed for their deployment flow. Deliverables: • Clean Python/Flask project with FER, OpenCV and TensorFlow wired together • Responsive single-page UI that streams the webcam feed and overlays the detected emotion as text • Read-me with setup, local run instructions and Render deployment steps If the above sounds straightforward, let’s get started. Skills: JavaScript, Python, CSS, Django, HTML, Git, OpenCV, Web Development, Flask, Deep Learning
Fixed budget:
12,500 - 37,500 INR
1 day ago
|
|||||
|
Computer Vision Project – Automated Product Detection and Image Analysis
Applied
|
not specified | 1 day ago |
Client Rank
- Excellent
$25 117 total spent
55 hires, 2 active
67 jobs posted
82% hire rate,
7 open job
35.42 /hr avg hourly rate paid
616 hours paid
Industry: Media & Entertainment
Individual client
Registered: Apr 15, 2025
Liepaja
15:24
5
|
||
|
We are looking to develop a computer vision solution for automated analysis of product and retail images.
The system should be able to process images captured in real-world environments, identify relevant products or objects, and organize the extracted visual information into a structured and usable format. Project objectives: Detect and recognize products or selected object categories within images Handle different camera angles, lighting conditions, image quality, and partial occlusion Extract useful visual data such as object location, quantity, category, or condition Support processing of multiple images and larger image datasets Provide results through a simple interface, dashboard, or API Build a foundation that can later be expanded with additional recognition and analytics features We are looking for an end-to-end project implementation, including solution architecture, model development or integration, backend processing, testing, and deployment. The initial phase may include a prototype or proof of concept, followed by further development based on the achieved results. Relevant experience with computer vision, object detection, image processing, and production-ready ML systems would be valuable.
Budget:
not specified
1 day ago
|
|||||
|
Computer Vision Project – Automated Product Detection and Image Analysis
Applied
|
not specified | 1 day ago |
Client Rank
- Excellent
$25 117 total spent
55 hires, 2 active
67 jobs posted
82% hire rate,
7 open job
35.42 /hr avg hourly rate paid
616 hours paid
Industry: Media & Entertainment
Individual client
Registered: Apr 15, 2025
Liepaja
15:24
5
|
||
|
We are looking to develop a computer vision solution for automated analysis of product and retail images.
The system should be able to process images captured in real-world environments, identify relevant products or objects, and organize the extracted visual information into a structured and usable format. Project objectives: Detect and recognize products or selected object categories within images Handle different camera angles, lighting conditions, image quality, and partial occlusion Extract useful visual data such as object location, quantity, category, or condition Support processing of multiple images and larger image datasets Provide results through a simple interface, dashboard, or API Build a foundation that can later be expanded with additional recognition and analytics features We are looking for an end-to-end project implementation, including solution architecture, model development or integration, backend processing, testing, and deployment. The initial phase may include a prototype or proof of concept, followed by further development based on the achieved results. Relevant experience with computer vision, object detection, image processing, and production-ready ML systems would be valuable.
Budget:
not specified
1 day ago
|
|||||
|
Real-Time Parking Vehicle Counter
Applied
|
~131 - 392 USD
|
1 day ago |
Client Rank
- not enough data
-
|
||
|
I need a lightweight, camera-agnostic application that taps into the IP feeds already installed at my parking lot and shows, live, how many cars and motorcycles have entered and exited. Accuracy and low latency are critical because the numbers will feed straight into our occupancy display as well as daily summaries we archive.
Here’s what I’m expecting: • Software (GUI or web dashboard) that connects to multiple RTSP or ONVIF streams and automatically detects entry and exit lines. • Separate, real-time tallies for cars and motorcycles, with the running balance always visible. • A simple way to reset or export counts (CSV/JSON) at the end of a chosen period. • Installation guide plus brief documentation so my team can add new cameras later. OpenCV, YOLOv8, TensorFlow, or a comparable computer-vision stack is fine as long as the final solution runs reliably on a mid-range Windows or Linux box without expensive GPUs. I’ll consider the project complete when I can point the finished program at two of our existing cameras, watch vehicles come and go, and see the live numbers update with at least 95 % accuracy verified over a one-hour test. Skills: PHP, Python, Software Architecture, C++ Programming, OpenCV, Computer Vision, AI Model Development, AI Integration
Fixed budget:
12,500 - 37,500 INR
1 day ago
|
|||||
|
Japanese OCR / Document AI SDK QA Engineer (Linux & Python)
Applied
|
$1,500
|
1 day ago |
Client Rank
- Good
$4 800 total spent
1 hires, 2 active
6 jobs posted
17% hire rate,
3 open job
Registered: Aug 9, 2025
KANAGAWA
21:24
4
|
||
|
PROJECT OVERVIEW
We are looking for an experienced QA engineer to evaluate a document extraction and OCR SDK for Japanese-language documents. Budget: USD 1,500 fixed price Duration: Approximately 10 business days of active work Target completion: August 31, 2026 Start: As soon as possible The SDK runs on Linux x64. The selected contractor will install the SDK, test it against Japanese invoices, contracts, forms, and scanned documents, measure extraction accuracy, document reproducible defects, and develop an automated regression test program. This project is focused on measuring and documenting the SDK's current quality. Improving the OCR engine to reach a specific accuracy target or making major changes to the SDK itself is not part of the scope. The SDK, license, English documentation, and available test materials will be provided after contractor selection and, if required, execution of an NDA. SCOPE LIMITS - Up to 50 document files, with a maximum of 150 pages in total - One agreed Linux x64 environment - Up to 20 Japanese UI screens and 30 pages of documentation for language review - One initial test cycle and one corrected-SDK retest - One agreed set of extraction fields - Waiting time for a corrected SDK is not included in the 10 active business days SCOPE OF WORK 1. Environment setup - Prepare or use a Linux x64 test environment - Install the SDK and required dependencies - Configure licensing or authentication - Confirm supported input and output formats - Run an initial smoke test - Review SDK errors and logs 2. Test planning - Classify document types and extraction fields - Define test cases and expected results - Define the ground-truth data format - Define accuracy metrics and defect severity levels - Confirm acceptance criteria before full testing 3. Test data and ground truth - Organize up to 50 approved documents / 150 pages total - Create or normalize expected results in JSON or CSV - Map each test document to its expected output - Ensure that personal and confidential information is handled appropriately Expected document types may include: - Japanese invoices, contracts, application forms, and business forms - Scanned, skewed, noisy, or low-resolution documents - Documents containing vertical Japanese text - Mixed kanji, hiragana, katakana, Latin characters, and numbers - Tables and line items - Stamps or seals - Dates, amounts, currencies, names, companies, and addresses - Multi-column or otherwise complex reading order Test data may consist of approved public data, synthetic data, or data supplied securely by the client. 4. OCR and extraction testing Evaluate items such as: - Company and personal names - Addresses, telephone numbers, and email addresses - Invoice, document, and contract numbers - Issue dates and payment due dates - Subtotals, tax, totals, and currencies - Product or service names, quantities, and unit prices - Line items and table structures - Stamps or seals, where supported - Text reading order Also record crashes, timeouts, encoding problems, and unexpected errors. 5. Accuracy evaluation Compare SDK output with ground truth and classify results as: - Exact match - Partial match - Missing extraction - Incorrect extraction - Incorrect field assignment - Character corruption - Broken table structure - Incorrect reading order Where appropriate, calculate field-level accuracy, document-level accuracy, precision, recall, F1 score, and character or word error rate. Reaching a specific accuracy level is not an acceptance requirement. 6. Defect investigation Each reported problem should include: - Issue summary and severity - Affected document and input conditions - Expected and actual results - Reproduction steps - Relevant SDK output and logs - Screenshot or other supporting evidence - Reproduction frequency - Suggested workaround or improvement, where possible 7. Automated regression test program Develop a reproducible test program, preferably in Python, that can: - Process all documents in a specified directory - Execute the SDK automatically - Save raw SDK output - Compare output with JSON or CSV ground truth - Detect differences and determine pass/fail status - Aggregate results by document and field - Export results in CSV and/or JSON - Save execution logs - Compare the current run with a previous run Another language may be used if required by the SDK interface. 8. Japanese UI and documentation review Review the available Japanese UI or Japanese-facing materials for: - Unnatural Japanese - Translation errors or inconsistent terminology - Buttons and error messages - Display problems in a Japanese environment - Missing or unclear setup instructions - Areas likely to confuse Japanese users If editable source files are unavailable, provide proposed corrections in the final report. 9. One regression retest If a corrected SDK is supplied within the agreed schedule, perform one retest to confirm: - Previously reported problems have been addressed - Existing document processing still works - Accuracy has not materially regressed - No new crashes or major errors have appeared DELIVERABLES - Test plan and test case list - Approved test data and structured ground-truth data - Regression test source code - Environment setup and execution instructions - Document-level and field-level accuracy results - Defect list - Reproduction steps, logs, screenshots, and supporting evidence - Japanese UI and documentation improvement proposals - Results of one corrected-SDK retest, if the SDK is supplied on schedule - Final report and handover materials Any restrictions on redistributing third-party or confidential test data must be documented. ACCEPTANCE CRITERIA The project will be accepted when: - The SDK can be executed in the agreed Linux x64 environment - Up to 50 agreed documents / 150 pages have been tested - Results are recorded for the agreed extraction fields - Reported defects contain evidence and reproduction instructions - The regression program can be rerun using the provided instructions - Source code and accuracy reports have been delivered - Japanese UI and documentation issues have been documented - One retest has been completed if the corrected SDK is provided within the agreed schedule - Final reporting and handover are complete Acceptance does not require the SDK to reach a specific accuracy level or for every reported defect to be fixed. OUT OF SCOPE - Major SDK source-code modifications - Development of a new OCR engine - AI model training or fine-tuning - Production integration - Building or operating commercial infrastructure - Testing beyond the agreed document/page limit - More than one corrected-SDK retest - Extensive UI or documentation rewriting outside the agreed limits - Ongoing production-data processing - Guaranteeing OCR accuracy Additional work will require a separate estimate and milestone. SECURITY REQUIREMENTS - Do not upload documents to external services without written approval - Do not reuse test data for another purpose - Protect personal and confidential information - Sign an NDA if required - Follow the agreed deletion procedure after completion - Do not disclose the SDK, source code, or test results to third parties REQUIRED QUALIFICATIONS - Professional or native-level Japanese - Experience testing OCR, document AI, or document extraction systems - Strong Python test-automation experience - Linux x64 development and troubleshooting skills - Experience working with JSON, CSV, APIs, logs, and command-line tools - Ability to build structured ground-truth datasets - Understanding of precision, recall, F1, and OCR error metrics - Clear written reporting in English - Ability to handle confidential materials securely Experience with Japanese invoices, contracts, table extraction, reading-order evaluation, or image preprocessing is a plus. PROPOSED MILESTONES 1. SDK setup, smoke test, test plan, and test-case design: USD 250 2. Ground truth, full testing, accuracy evaluation, defect evidence, and regression program: USD 950 3. One corrected-SDK retest, final report, and handover: USD 300 Total: USD 1,500 PLEASE INCLUDE IN YOUR PROPOSAL 1. Whether you can complete the project 2. Your earliest available start date 3. Whether you can complete approximately 10 business days of active work by August 31, 2026 4. Number of team members and their roles 5. Relevant OCR, document AI, Japanese-language testing, and automation experience 6. Estimated effort for each project phase 7. Proposed technologies and test environment 8. What you require from us before starting 9. Confirmation that you accept the USD 1,500 fixed budget 10. Assumptions, possible additional costs, and current questions Please briefly describe one relevant OCR or document-processing project and your specific contribution.
Fixed budget:
1,500 USD
1 day ago
|
|||||
|
C/C++ Developer for Orbbec K4A Wrapper
Applied
|
not specified | 1 day ago |
Client Rank
- Risky
1 open job
South Korea
21:24
1
|
||
|
We are seeking an experienced C/C++ developer familiar with Azure Kinect, Orbbec Femto Bolt, RGB-D cameras, Windows native DLLs, or camera SDKs. Our application uses the Orbbec K4A Wrapper, k4a.dll, and k4abt.dll for depth capture and body tracking. We need to address two main issues: application crashes in k4abt.dll and RGB stream freezing upon depth stream restart. The developer should reproduce and analyze the crash, review memory release and multithreading, and modify the wrapper or k4a.dll if necessary. Required skills include C/C++, Visual Studio, CMake, Windows native debugging, and camera SDK integration. Experience with Azure Kinect, Orbbec SDK, WinDbg, ONNX Runtime, CUDA, or DirectML is preferred.
Client's questions:
Budget:
not specified
1 day ago
|
|||||
|
Senior Full-Stack Developer for Secure AI-Assisted OCR and Document Redaction Platform
Applied
|
$1,000
|
1 day ago |
Client Rank
- Good
$3 218 total spent
18 hires, 1 active
64 jobs posted
28% hire rate,
1 open job
32.16 /hr avg hourly rate paid
45 hours paid
Industry: Education
Individual client
Registered: Oct 26, 2017
Las Vegas
09:24
4
|
||
|
I revised the description you posted to preserve the core platform, security, permanent-redaction, ownership, and discovery requirements while making it more suitable for a worldwide Upwork posting.
This version reflects a global talent search , a fixed-price discovery phase , separate milestones, and a detailed requirements package shared only with shortlisted applicants. Upwork permits global job posts, fixed-price milestones, and separate NDAs; pre-contract communication should remain on Upwork. # Senior Full-Stack Developer for Secure AI-Assisted Document Redaction Platform ## Project Overview We are seeking an experienced senior full-stack developer, technical lead, or small coordinated development team to help design and eventually build a secure, AI-assisted document redaction and records-processing platform. Applicants may be located anywhere. However, all proposed team members, developers, specialists, and subcontractors must be identified before receiving access to the project. This is not a request for: * A basic website * A general chatbot * A simple PDF editor * A basic file-storage application * A public AI wrapper * An automated redaction tool without human review * An application that only places black boxes over visible text The platform will support a managed document-processing operation that receives electronic archives, scanned records, historical documents, and document backlogs from government agencies, regulated organizations, legal offices, businesses, and other clients. Authorized personnel will use the platform to: * Receive client documents securely * Organize records by client, project, batch, file, and page * Perform OCR and document-image processing * Classify documents and pages * Detect potentially protected or confidential information * Present suggested redactions to trained human reviewers * Require human validation of redactions * Conduct secondary quality-control review * Permanently redact and sanitize approved documents * Verify that removed information cannot be recovered * Generate reports, manifests, metadata, and delivery packages * Return completed files securely * Track retention and secure deletion * Produce audit reports and deletion certificates Accuracy, confidentiality, security, human validation, secondary quality control, permanent redaction, records integrity, and traceability are essential requirements. ## Location and Communication Requirements This is a worldwide opportunity. Applicants must: * Identify the country and time zone of every person who will work on the project * Disclose whether the work will be performed by an individual, agency, employees, partners, or subcontractors * Provide at least two hours of communication overlap with Pacific Time * Communicate clearly in English * Attend scheduled video meetings when required * Provide regular written progress updates * Obtain written approval before adding or replacing team members Development may be performed from any approved location. However, live client records and production document processing will remain in a company-controlled environment located in the United States. ## Initial Contract: Paid Discovery Phase The first contract will be a fixed-price discovery, requirements, architecture, and technical-planning engagement. The selected developer will not immediately build the complete production platform. The discovery phase will determine: * Functional requirements * Security requirements * User roles and permissions * Operational workflows * Recommended system architecture * Recommended technology stack * Database and file-storage design * OCR and image-processing approach * AI-assisted detection approach * Human-review workflow * Secondary quality-control workflow * Permanent-redaction methodology * Document-sanitization methodology * Audit-logging requirements * Reporting requirements * Development and production separation * Hosting and deployment architecture * Third-party tools and services * Licensing and recurring costs * Technical risks * Security risks * Proof-of-concept scope * Minimum viable product scope * Development milestones * Estimated schedule * Estimated development and operating costs After successful completion of discovery, the selected developer may be considered for additional milestones involving a technical proof of concept, minimum viable product, security testing, deployment, documentation, maintenance, and support. ## Hosting and Deployment The company does not currently plan to purchase physical server equipment. During discovery, the selected developer must recommend an appropriate company-controlled cloud, private-cloud, dedicated-hosting, or hybrid architecture based on: * Document volume * Page volume * File sizes * OCR requirements * AI-processing requirements * Storage requirements * Security requirements * Backup and disaster-recovery requirements * Client requirements * Expected future growth * Estimated operating costs Development should initially use a company-controlled cloud environment with separate development, testing, staging, and production configurations. All hosting accounts, domains, databases, storage resources, administrator accounts, credentials, encryption keys, source-code repositories, and production configurations must remain under company control. ## Required Platform Capabilities ### Secure Document Intake The platform should support: * Secure client and employee accounts * Individual and bulk document uploads * Large files and large document batches * SFTP or another secure transfer method * Upload progress and status reporting * File-format validation * Malware and virus scanning * Duplicate-file detection * File-integrity verification * Intake manifests * Project and batch identification * File and page inventories * Chain-of-custody tracking * Failed-upload reporting * Exception reporting * Configurable retention requirements ### Supported File Formats The system should support or be designed to support: * Searchable PDF * PDF/A * TIFF * JPEG * PNG * Microsoft Word files * OCR text * CSV * XML * JSON * Metadata files * Client-specific index and import files The architecture must allow additional document and output formats to be added later. ### OCR and Document-Image Processing Required or anticipated functions include: * OCR for scanned and image-based records * Preservation of page, line, word, and coordinate information * Page-orientation detection * Rotation and deskewing * Noise removal * Image enhancement * Blank-page detection * Searchable-text creation * OCR confidence scoring * Identification of unreadable or low-confidence pages * Manual OCR correction * Document classification * Page classification * Document separation * Document assembly * Batch processing * Reprocessing of failed or rejected files * Processing-status tracking Applicants should explain whether they recommend established OCR products, open-source OCR tools, private services, locally hosted technology, custom models, or a combination. ### AI-Assisted Protected-Information Detection The platform should assist trained reviewers with identifying protected, confidential, personal, or client-defined information, including: * Social Security numbers * Tax-identification numbers * Dates of birth * Driver’s license numbers * State-identification numbers * Passport numbers * Bank-account numbers * Credit-card information * Medical or health information * Signatures * Email addresses * Telephone numbers * Home addresses * Names of protected individuals * Information concerning minors * Legal case information * Property-record information * Client-defined names, words, phrases, patterns, fields, or page areas Detection methods may include: * Regular expressions * Pattern matching * Named-entity recognition * OCR coordinates * Document classification * Machine-learning models * Private AI services * Locally hosted models * Client-specific rules * Manual reviewer selections AI findings will be recommendations only. The system must not independently approve or finalize redactions. ### Configurable Client and Project Rules Administrators should be able to configure separate requirements based on: * Client * Agency * Project * Jurisdiction * Document type * Record series * Confidentiality category * Redaction category * Required output format * Quality-control level * Retention period * Delivery requirements The system should record which rule set and rule version were applied to each file. ## Human Review and Quality Control The platform must include a complete human-in-the-loop review process. Primary reviewers must be able to: * View the original document * View OCR text * Review AI-suggested redactions * Accept a suggested redaction * Reject a suggested redaction * Correct a suggested redaction * Resize or reposition a redaction area * Add a missed redaction * Assign a redaction reason or category * Add reviewer notes * Flag uncertain information * Escalate a document * Submit completed work for quality control Secondary quality-control reviewers must be able to: * Review the original file * Review proposed and approved redactions * Examine primary-review decisions * Approve completed work * Reject completed work * Return work for correction * Add quality-control findings * Escalate unresolved issues * Provide final approval Every action must be associated with the responsible user, date, time, project, batch, file, page, action, and result. The platform should also support: * Reviewer assignments * Supervisor review * Rework queues * Exception queues * Random quality-control sampling * Full quality-control review when required * Error categories * Corrective-action tracking * Reviewer accuracy reports * Productivity reports * Quality trends * Final completion approval ## Permanent Redaction and Document Sanitization The completed system must do more than place a visible rectangle or black box over information. Final processing must permanently remove or sanitize protected information from: * Visible page content * Underlying text * OCR text layers * Hidden objects * Hidden layers * Comments * Annotations * Form fields * Embedded attachments * Scripts and active content * Revision information * Document properties * Unapproved metadata * Thumbnail images * Temporary working files * Intermediate processing files * Other recoverable content The system should verify that protected information cannot be recovered by: * Copying and pasting * Selecting hidden text * Searching the completed file * Removing a visual overlay * Extracting the OCR text layer * Inspecting annotations * Opening embedded files * Reviewing metadata * Examining temporary or intermediate outputs The system must record redaction and sanitization validation results in the document’s audit history. ## Security and Data Boundary This engagement concerns software design and development. It does not include outsourced review or processing of live client records. The following requirements are mandatory: * Development and testing must use synthetic, simulated, or properly de-identified documents. * Developers will not have routine or unrestricted access to live client records. * Live records will remain in a company-controlled production environment. * Production document processing will occur in the United States. * Production databases, credentials, administrator accounts, encryption keys, and security configurations will remain under company control. * Development, testing, staging, and production environments must be separated. * Production records may not be copied into development or testing. * Client files may not be retained or used for demonstrations. * Project information may not be submitted to public consumer AI tools. * Client records may not be used to train public, private, commercial, personal, or developer-owned AI models. * Client information may not be retained by third-party services without prior written approval. * No work may be subcontracted without written approval. * No unidentified person may access the project. * Any exceptional production access must be approved, limited, time-restricted, monitored, logged, and capable of immediate termination. Applicants must identify every proposed external: * OCR service * AI service or model * PDF-processing service * Storage provider * Hosting provider * Logging or monitoring service * File-transfer service * Security service * Third-party library * Licensed software product The applicant must disclose the provider’s purpose, data flow, retention practices, licensing terms, and recurring costs. ## Application Security Requirements Required or anticipated controls include: * Role-based access * Least-privilege permissions * Multifactor authentication * Secure password controls * Encryption in transit * Encryption at rest * Secure secrets management * Secure encryption-key management * Session timeout controls * Account lockout protections * User activation and deactivation * Project-level access restrictions * Separation of client projects * Download restrictions * Access expiration * Administrative approval * Detailed audit logs * Security-event logging * Secure APIs * Input validation * File-integrity controls * Malware scanning * Backup and recovery * Retention controls * Secure deletion * Dependency scanning * Vulnerability scanning * Automated testing * Secure error handling * Production logging and monitoring Development should use recognized secure software-development practices capable of aligning with the NIST Secure Software Development Framework and using the OWASP Application Security Verification Standard as a basis for web-application security verification. ([NIST Computer Security Resource Center][2]) ## Audit Logging The platform should maintain detailed, meaningful, and tamper-resistant records of activities such as: * Login attempts * Authentication failures * Account changes * Permission changes * Document uploads * File validation * Document viewing * Reviewer assignments * Redaction decisions * Quality-control decisions * Administrative actions * File exports * Downloads * Deliveries * Retention changes * Deletion events * Security alerts * System errors Audit events should identify the user, date, time, affected client, project, batch, document, page, action, and outcome. ## Required Output Capabilities Depending on client requirements, the platform should be capable of producing: * Permanently redacted TIFF files * Searchable redacted PDFs * PDF/A files * OCR text files * Metadata files * CSV index files * XML index files * JSON index files * Client-specific import files * Batch manifests * File inventories * Exception reports * Redaction reports * Quality-control reports * Audit reports * Processing statistics * Chain-of-custody records * Secure-delivery confirmations * Retention reports * Deletion certificates ## Reporting and Dashboards The system should provide reports and dashboards showing: * Files and pages received * Files and pages processed * Processing status * OCR completion * OCR confidence * Documents awaiting review * Redactions proposed * Redactions accepted * Redactions rejected * Redactions corrected * Redactions manually added * Reviewer assignments * Reviewer productivity * Quality-control findings * Rework requirements * Error rates * Exception rates * Project completion percentage * Delivery status * Retention status * Deletion status * Estimated processing charges * Actual processing charges ## Future Integrations The architecture should permit future integration with: * Document-management systems * Records-management systems * Government records systems * Archival systems * Secure SFTP servers * Identity-management systems * Cloud-storage environments * Billing and accounting systems * Client databases * Records indexes * APIs * Secure web services The initial version does not need every future integration, but the architecture must support expansion. ## Preferred Experience Applicants should demonstrate relevant experience with several of the following: * Full-stack application development * Secure web-application development * Python * FastAPI or Django * React or Next.js * TypeScript * PostgreSQL * Redis * Background-processing queues * OCR * Computer vision * OpenCV * PDF processing * TIFF processing * Document classification * Named-entity recognition * Private AI models * Locally hosted AI models * Secure API development * Role-based authorization * Multifactor authentication * Audit logging * Secure file transfer * Docker * Cloud deployment * On-premises deployment * Automated testing * Vulnerability testing * Technical documentation General website, chatbot, or basic AI-wrapper experience alone is insufficient. ## Discovery-Phase Deliverables The first paid engagement should produce: 1. Functional-requirements specification 2. Security-requirements specification 3. User-role and permission matrix 4. Operational workflow diagrams 5. System-architecture diagram 6. Data-flow diagram 7. Preliminary database design 8. File-storage and processing design 9. OCR and document-processing recommendation 10. AI and protected-information detection recommendation 11. Permanent-redaction and sanitization design 12. Human-review and quality-control design 13. Audit-logging design 14. Environment-separation design 15. Hosting and deployment recommendation 16. Third-party technology list 17. Licensing and recurring-cost schedule 18. Technical-risk assessment 19. Security-risk assessment 20. Defined technical proof-of-concept scope 21. Defined minimum viable product scope 22. Milestone-based implementation plan 23. Estimated development schedule 24. Estimated proof-of-concept and MVP costs 25. Estimated ongoing hosting, maintenance, licensing, and support costs ## Potential Technical Proof of Concept Using synthetic documents, a later proof-of-concept milestone should demonstrate: * Secure document upload * OCR processing * Preservation of text coordinates * Identification of selected protected information * Human review of suggested redactions * Acceptance of a suggested redaction * Rejection of a suggested redaction * Correction of a suggested redaction * Manual addition of a missed redaction * Secondary quality-control review * Permanent redaction * Metadata sanitization * Redaction validation * Audit logging * Completed-file export ## Source-Code Ownership and Documentation The following requirements are mandatory: * The company must own all paid custom work product. * Source code must be stored in a private company-controlled repository. * Work must be committed regularly during development. * Completed source code may not be withheld until the end of the engagement. * Code must be readable, organized, tested, and maintainable. * Credentials and encryption keys may not be embedded in source code. * All open-source and third-party components must be disclosed. * All licensing, hosting, API, AI, OCR, maintenance, and recurring costs must be disclosed. * The developer may not reuse or resell confidential project materials. * The project may not be published in a portfolio without written approval. * Work may not be subcontracted without written approval. * Architecture, database, APIs, installation, configuration, deployment, security, administration, and maintenance procedures must be documented. * The completed system must be transferable to another qualified developer without requiring a complete rebuild. The selected applicant may be required to sign: * A nondisclosure agreement * A development or independent-contractor agreement * An intellectual-property and work-product assignment * A data-security and access agreement * A subcontractor disclosure ## Proposal Instructions Begin your proposal with: **SECURE DOCUMENT PLATFORM** Then address the following: 1. Identify your location, time zone, and available Pacific Time overlap. 2. State whether you personally perform the work. 3. Identify every person who would have access to the project. 4. Disclose all employees, partners, agencies, or subcontractors who may participate. 5. Describe your experience with OCR and scanned-document processing. 6. Describe your experience with PDF and TIFF processing. 7. Explain your experience with permanent redaction and metadata sanitization. 8. Describe a human-review or quality-control workflow you developed. 9. Explain how you would separate development, testing, staging, and production. 10. Explain how you would prevent unauthorized access to live production records. 11. Provide two relevant project examples and explain which portions you personally completed. 12. Identify your preliminary recommended technology stack. 13. Identify any OCR, AI, PDF, storage, hosting, or security products you would consider. 14. Confirm that source code can remain in a private repository controlled by the company. 15. Confirm that project information and records will not be used for AI training. 16. Confirm that you will not subcontract the work without written approval. 17. Provide a fixed-price estimate and timeline for the discovery phase. 18. Provide a preliminary cost range for the technical proof of concept. 19. State your weekly availability. 20. Describe your proposed milestone and payment structure. Please provide specific responses. Generic proposals will not be considered. ## Contract Structure The initial contract will be a fixed-price discovery phase with defined deliverables, deadlines, acceptance criteria, and milestone payments. Potential additional milestones may include: 1. Technical proof of concept 2. Core platform development 3. Human-review and quality-control functions 4. Reporting and administration 5. Security testing 6. Production preparation 7. Deployment and documentation 8. Maintenance and support Fixed-price milestones allow the project to be divided into defined portions of work with separate deliverables and payments. ([Upwork Support][3]) ## Selection and Next Steps Shortlisted applicants may be invited to: * Participate in an Upwork video interview * Explain a relevant OCR, document-processing, secure SaaS, or redaction project * Complete a small paid technical evaluation * Review and sign required agreements * Review the complete project-requirements package * Submit a final fixed-price discovery proposal A detailed project-requirements package will be provided through Upwork to shortlisted applicants. An NDA may be required before nonpublic business, workflow, architecture, or security information is shared. Cost is important, but the lowest proposal will not automatically be selected. Selection will consider: * Relevant document-processing experience * Permanent-redaction knowledge * Security awareness * Full-stack technical ability * Communication * Documentation * Reliability * Cost * Availability * Ability to preserve the platform’s purpose and requirements All pre-contract communication, interviews, file sharing, and negotiations must remain within Upwork. Do not include an outside email address, telephone number, or meeting link in the public posting. Client's questions:
Fixed budget:
1,000 USD
1 day ago
|
|||||
|
Senior Full-Stack Engineer (Python/Node.js) – AI Media & Automation Engine (MVP)
Applied
|
$3,000
|
1 day ago |
Client Rank
- Medium
1 open job
15:24
3
|
||
|
We are a stealth-mode startup building an innovative AI automation platform for the e-commerce and media technology market.
We are looking for a Senior Full-Stack Engineer to build an isolated media processing and AI data extraction engine as part of our core MVP. WHAT WE ARE BUILDING (SCOPE): You will be responsible for developing a standalone Python or Node.js module that accepts short user video inputs, processes frames using computer vision (FFmpeg/OpenCV), transcribes audio (Whisper API), and utilizes Multimodal AI models (OpenAI GPT-4o / Gemini 1.5 Flash) to generate structured JSON data outputs. WHAT IS ALREADY PREPARED: - Fully structured Technical Specification & PRD hosted on GitHub. - Clear API schemas, JSON outputs, and architecture guidelines. - Standardized IP Assignment & Confidentiality Agreement. REQUIRED TECHNICAL EXPERTISE: - Backend: Python (FastAPI / Django) OR Node.js / TypeScript. - AI & Computer Vision: OpenAI API (GPT-4o, Whisper), Gemini API, FFmpeg / OpenCV for frame extraction. - Browser Automation (Bonus): Hands-on experience with Puppeteer, Playwright, or Selenium. - APIs & Cloud: RESTful APIs, PostgreSQL/Redis, Docker. PROJECT TERMS & PROCESS: - Project Type: Fixed-Price MVP development (with potential long-term contract / Lead Dev role). - Estimated Timeline: 3 to 4 weeks. - Workflow: A 1-week paid test sprint will be conducted with shortlisted candidates. - Confidentiality: NDA and IP Assignment Agreement MUST be signed prior to full PRD and repository access. HOW TO APPLY: 1. Briefly describe your experience with AI API integrations, video frame extraction, or web automation. 2. Provide 1–2 links to relevant past projects, live demos, or your GitHub profile. 3. Confirm your availability to start within the next 3–5 days and your willingness to sign an NDA/IP agreement. Client's questions:
Fixed budget:
3,000 USD
1 day ago
|
|||||
|
Computer Vision & Deep Learning Engineer (PyTorch/OpenCV)
Applied
|
$40 - $60
/ hr
|
1 day ago |
Client Rank
- Risky
1 open job
Registered: Apr 27, 2024
Bauru
09:24
1
|
||
|
We are a specialized AI/ML engineering team working on a high volume of complex projects across diverse industries (including Health Tech, AgTech, and Industrial Automation). As our project pipeline grows, we are looking to expand our core team with a highly skilled and hands-on Computer Vision Engineer.
This is an ongoing, long-term opportunity to join a collaborative environment and work alongside other senior AI professionals to tackle real-world vision problems, from 2D object detection to 3D image segmentation and feature extraction. Core Responsibilities: - Collaborate with our existing engineering team on the design and implementation of computer vision pipelines (Object Detection, Semantic Classification, Image Segmentation). - Train, fine-tune, and evaluate deep learning models using PyTorch and TensorFlow. - Perform data preprocessing, augmentation, and handling of large-scale image/video datasets. - Optimize models for inference and assist in integrating computer vision APIs into our broader software architecture. - Write clean, well-documented, and production-ready Python code while participating in code reviews with the team. Mandatory Skills & Expertise: - Strong proficiency in Python and core computer vision libraries (OpenCV). - Proven hands-on experience with PyTorch (preferred) or TensorFlow. - Solid background in implementing architectures like YOLO, SSD, or similar for object detection/segmentation. - Experience with deep learning model evaluation and performance tuning. Nice to Have: - Experience with 3D Reconstruction, Stereo Matching, or Multi-view Geometry. - Knowledge of C++ and CUDA optimization. Familiarity with SLAM, Robot Operating System (ROS), or autonomous vehicle applications.
Hourly rate:
40 - 60 USD
1 day ago
|
|||||
|
Project: Grandstream UCM6301 – Central Phonebook & Wave App Setup (Fixed Price)
Applied
|
not specified | 1 day ago |
Client Rank
- Risky
1 jobs posted
1 open job
Registered: Jul 29, 2026
14:24
1
|
||
|
Hello,
We are looking for an experienced Grandstream specialist to help us complete the setup of our phone system. Our current setup PBX Grandstream UCM6301 Phones 2 × Grandstream GRP2613 2 × Grandstream GRP2624 1 × Yealink SIP-T58W The telephony itself is already fully configured and working correctly. We are not looking for help with the PBX configuration or SIP extensions. What we need 1. Central Phonebook Configure a centralized phonebook on the UCM6301. We want to be able to add, edit, and delete contacts from one central location. All Grandstream phones should automatically synchronize with the central phonebook. The phonebook search function should work on all devices. 2. Grandstream Wave App Set up the free Grandstream Wave app on all of our smartphones. Configure each user with their extension. Ensure the app is fully functional for making and receiving calls through the UCM6301. Remote Access Remote access can be provided. Pricing We are only looking for a fixed-price project. We are not interested in hourly billing. If you have experience with Grandstream UCM systems, please let us know: Your experience with similar projects. Your fixed price for completing this project. When you would be available to start. We look forward to hearing from you.
Budget:
not specified
1 day ago
|
|||||
|
UAV Facial Detection Simulation Development
Applied
|
~27 - 332 USD
|
1 day ago |
Client Rank
- not enough data
-
|
||
|
I need a developer to create a simulation-based UAV/drone facial detection and tracking system. I am not building or flying a real drone. The aim is to create a software simulation where a virtual drone camera moves through an environment and detects/tracks human faces from the camera feed.
The project should be easy to understand, realistic to complete quickly, and use free or low-cost software. The preferred approach is: Option 1 — Best/easiest preferred route: Use Webots for the drone/3D simulation, with a virtual camera attached to a simulated UAV. Then use Python + OpenCV + MediaPipe or another lightweight face detector to detect and track faces from the simulated camera feed. Option 2 — Acceptable simpler route: If Webots is too difficult, create a simplified Python/OpenCV simulation using UAV-style video footage or a custom 3D/animated environment that simulates a drone camera moving above/around people. It still needs to clearly look like a drone-camera facial detection/tracking demo. The system should include: • A simulated UAV/drone camera view • Face detection from the simulated camera feed • Face tracking across frames • Bounding boxes around detected faces • Basic tracking IDs if possible • Output video showing the detection/tracking results • Performance results such as FPS, number of faces detected, detection confidence if available, and tracking stability • Screenshots of the simulation and detection output • Clear setup instructions so I can run it myself The final delivery must include: 1. Working source code o Python scripts o Any Webots world/project files if used o Requirements file, e.g. requirements.txt o Clear folder structure 2. Step-by-step documentation o Exactly what software was installed o Which websites were used o How the simulation was created o How the virtual drone/camera was set up o How the face detection/tracking code works o How to run everything from start to finish o Include screenshots at each major stage 3. Final outputs o Annotated output video showing faces being detected/tracked o Screenshots of the running simulation o Screenshots of the Python/OpenCV detection output o CSV or simple results file showing FPS/detections per frame if possible o A short explanation of limitations and possible improvements Suggested software/websites: • Webots: for the drone/robot simulation • Python: main programming language • OpenCV: video processing and drawing bounding boxes • MediaPipe Face Detection or OpenCV YuNet: face detection • NumPy / Pandas / Matplotlib: for results and graphs • VS Code: code editor • GitHub or ZIP folder: for final code delivery Important requirements: • Please keep the project simple and practical. • Do not make it depend on expensive hardware. • Do not use paid APIs. • Do not require a real drone. • Use clear comments in the code. • The simulation and code must be easy for a beginner/intermediate student to run. • The project should be completed in a way that can be explained clearly with screenshots and a written technical breakdown. I need someone who can provide both the working simulation/code and a detailed explanation of exactly what they did, including screenshots and setup steps. Skills: Python, Machine Learning (ML), Robotics, Artificial Intelligence, OpenCV, NumPy, Video Processing, Software Documentation, Computer Vision, Simulation
Fixed budget:
20 - 250 GBP
1 day ago
|
|||||
|
Computer Vision Developer - Object Detection & Tracking on Video (Python)
Applied
|
$4 - $6
/ hr
|
1 day ago |
Client Rank
- Medium
$200 total spent
1 hires
1 jobs posted
100% hire rate,
2 open job
Registered: Jul 6, 2026
yerevan
16:24
3
|
||
|
I'm building a small sports analytics tool and need help with the video processing part.
What I need: - Detect a ball and players on recorded match footage; - Track them across frames and keep consistent IDs; - Output detections as JSON (coordinates, frame number, object ID). The footage is fixed-camera, 1080p. I already have sample videos to work with. Models can be off-the-shelf - YOLO, ByteTrack, whatever you're comfortable with. No need to train from scratch at this stage. Looking for someone who can start with a small piece first so we can see how we work together, then continue if it goes well. Long-term collaboration possible if the first part goes smoothly. Please tell me in your reply: - What you'd use for detection and tracking, and why? - A project where you did something similar (link if possible)? - How many hours per week you're available?
Hourly rate:
4 - 6 USD
1 day ago
|
|||||
|
Point Cloud & Mesh Reconstruction Expert (Open3D, RGB-D, LiDAR)
Applied
|
$2,000
|
2 days ago |
Client Rank
- Risky
1 open job
Hyderabad
17:54
1
|
||
|
We are looking for an experienced 3D Computer Vision / Point Cloud Developer to develop an end-to-end 3D reconstruction pipeline using Open3D, RGB-D, and LiDAR data.
The project involves processing raw RGB images, LiDAR depth maps, camera poses, calibration data, and related metadata to generate an accurate and dense 3D point cloud and high-quality mesh. The reconstructed model will be used for industrial inspection and comparison against reference CAD models, with a target geometric accuracy of approximately ±10 mm. What We Need The selected freelancer will be responsible for developing a complete processing pipeline that includes: RGB-D and LiDAR data processing Multi-frame point cloud registration and alignment ICP and global registration Noise, outlier, and drift removal Pose graph optimization Point cloud filtering and denoising Voxel downsampling and normal estimation Surface reconstruction and mesh generation Camera calibration and coordinate transformations Optimization of reconstruction accuracy and performance Clean, modular, and well-documented Python code Required Experience Open3D experience is mandatory. Candidates should have proven experience with: Open3D LiDAR point cloud processing RGB-D processing Point cloud registration ICP / Global Registration Pose Graph Optimization Surface reconstruction Mesh generation Camera calibration 3D geometry and linear algebra Python-based 3D processing Experience with PCL, ROS/ROS2, SLAM, SfM, OpenCV, CUDA/GPU processing, digital twins, robotics, metrology, or industrial inspection will be an advantage. Expected Deliverables The freelancer will deliver: Complete Open3D processing pipeline Well-structured Python source code Technical documentation Sample reconstructed point cloud and mesh Recommendations for improving reconstruction accuracy and processing performance
Fixed budget:
2,000 USD
2 days ago
|
|||||
|
Senior C++ / Android / iOS Engineer for Advanced AI Media Processing Platform
Applied
|
$3,000
|
2 days ago |
Client Rank
- Medium
$300 total spent
5 hires, 1 active
12 jobs posted
42% hire rate,
2 open job
Registered: Mar 15, 2023
Abu Dhabi
12:24
3
|
||
|
Senior Native Mobile Engineer (C++ / Android / iOS)
We are looking for an experienced senior software engineer (or a very small senior engineering team) to build a high-performance cross-platform mobile application. This project involves advanced native mobile development, media processing, computer vision, GPU acceleration, and on-device AI. The successful candidate will work from a detailed engineering specification. The complete project documentation will only be shared with shortlisted candidates after execution of a Mutual Non-Disclosure Agreement (NDA). Required Technical Experience Applicants should have strong experience in most of the following: Modern C++ (C++17/C++20) Kotlin Swift Android NDK JNI Jetpack Compose SwiftUI CameraX AVFoundation FFmpeg OpenCV MediaCodec Metal Vulkan Core ML Native mobile optimization GPU acceleration Cross-platform architecture Performance optimization Memory optimization Thermal optimization Battery optimization Production software development Responsibilities The selected engineer will be responsible for building a production-quality native mobile application with: Shared C++ core Native Android implementation Native iOS implementation Advanced media processing pipeline AI model integration Performance optimization Cross-platform architecture Testing Documentation Production-ready source code Deliverables The successful candidate must deliver: Complete source code Full Git repository Android project iOS project Technical documentation Build instructions Test suite Production-ready release Important This is a fixed-price project ($3,000). Expected duration: 10–12 weeks. Only candidates with strong, verifiable production experience should apply. The project specification will be released only after NDA execution. Please answer the following questions in your proposal: How many years of professional C++ experience do you have? How many native Android applications have you shipped? How many native iOS applications have you shipped? Have you built cross-platform applications using a shared C++ core? Describe your experience with FFmpeg and OpenCV. Describe your experience with MediaCodec and AVFoundation. Describe your experience with Metal, Vulkan, or GPU optimization. Have you integrated AI models into mobile applications? If yes, briefly describe the project. Please provide links to your most relevant applications. What was the most technically challenging mobile project you have completed, and what was your exact contribution? Client's questions:
Fixed budget:
3,000 USD
2 days ago
|
|||||
|
Multi Camera Activity Analysis
Applied
|
$30 - $250
|
2 days ago |
Client Rank
- not enough data
-
|
||
|
I’m building a real-time activity analysis pipeline that ingests live streams from well over ten IP cameras and flags everything that matters to my operations team. The focus is threefold: accurate people counting, reliable intrusion detection, and fluid crowd-movement analysis.
At the core I expect a YOLO-based model (v5, v7 or v8—you can advise) running through OpenCV that can scale horizontally as additional RTSP streams come online. Low-latency processing, smart use of GPU resources, and clean separation between detection and business-logic layers are crucial because the system will eventually tie into an existing alert dashboard. Deliverables • End-to-end Python (or C++) code that connects to each camera, performs the detections described above, and outputs structured JSON or MQTT topics I can consume in my backend. • Simple CLI or minimal web UI to visual-debug results from any selected camera feed. • Setup guide covering environment, dependencies, and sample configuration for adding new cameras. • Short performance report demonstrating FPS, detection accuracy, and resource usage with at least ten concurrent streams. If you have prior benchmarks or repos that show similar high-camera-count deployments, that will help us get started faster. Skills: Python, Machine Learning (ML), C++ Programming, OpenCL, OpenCV, Computer Vision, Deep Learning, YOLO
Fixed budget:
30 - 250 USD
2 days ago
|
|||||
|
Automated video analysis
Applied
|
$100
|
2 days ago |
Client Rank
- Risky
1 open job
09:24
1
|
||
|
Title: Python Developer Needed to Automatically Catalogue Comics from Video into Excel
I have overhead iPhone videos showing comic books being placed on a table one at a time. I need a simple Python-based tool that processes the video and creates an Excel spreadsheet containing one row for each comic. The tool should: 1. Detect each comic shown in the video. 2. Capture the clearest frame of each comic cover. 3. Crop and straighten the cover image. 4. Identify the comic as accurately as possible using the cover, visible text, barcode, OCR, online comic databases, or an existing image-recognition API. 5. Create an Excel spreadsheet with: * Embedded cover image * Series title * Issue number * Publisher * Publication year, when available * Confidence or review status * Image filename * Notes column 6. Clearly mark comics that cannot be identified confidently so I can review them manually. The camera position, lighting, and background will remain consistent. Comics will be shown individually with a brief empty-table gap between them. Requirements: * Runs locally on Windows * Full source code included * Simple setup instructions * May use existing open-source projects and APIs * A basic script is acceptable; I do not need a polished application or custom interface This is a proof-of-concept project. Please explain briefly how you would approach it and provide examples of relevant Python, OpenCV, OCR, image-processing, or inventory automation work.
Fixed budget:
100 USD
2 days ago
|
|||||
|
AI-Driven PDF Data Extraction Solution
Applied
|
~131 - 392 USD
|
2 days ago |
Client Rank
- not enough data
-
|
||
|
AI-Powered PDF Data Extraction & OCR Automation
Need to extract structured data from complex PDF documents with high accuracy? I can build custom Python-based automation solutions for converting scanned or digital PDFs into Excel, CSV, JSON, or databases. What I Can Build PDF to Excel / CSV Automation OCR for Scanned Documents AI-powered Document Data Extraction Bulk PDF Processing Multi-language Document Support (English, Hindi, Telugu and more) Duplicate Detection & Data Validation Intelligent Error Detection & Record Verification Offline or Low-Cost Processing Solutions Technologies Python OpenCV Tesseract OCR / EasyOCR PyMuPDF pdfplumber Pandas OpenPyXL AI / Computer Vision Excel & CSV Automation Suitable Projects Voter List Extraction Invoice Processing Bank Statement Extraction Government Documents Forms & Applications Survey Reports ID Cards & Certificates Property Records Legal Documents Custom PDF Data Extraction Workflows What You Get Clean and accurate structured data Reusable automation scripts Source code included Fast bulk processing Custom validation rules Well-documented solution If you're looking for a reliable Python developer for PDF automation, OCR, or AI-based document extraction, feel free to get in touch. I'll build a solution tailored to your workflow. Skills: Python, Data Processing, Web Scraping, Software Architecture, OpenCV, Data Extraction, Data Analysis, Pandas, AI Development, OCR Automation
Fixed budget:
12,500 - 37,500 INR
2 days ago
|
|||||
|
Python-Based Document Scanner Needed
Applied
|
$250 - $750
|
2 days ago |
Client Rank
- not enough data
-
|
||
|
# Python Developer Required – South African Driver’s Licence and Vehicle Licence Disc Decoder
We are looking for an experienced Python developer to create a solution that can scan and decode: 1. South African driver’s licences. 2. South African motor vehicle licence discs. The solution will be incorporated into our existing web-based software and must be usable from a mobile device. Users should be able to use their mobile phone camera to scan either document, after which the decoded information must be returned in a structured JSON format. ## Project Requirements The solution must: * Scan a South African driver’s licence using a mobile device camera. * Read and decode the barcode on the driver’s licence. * Extract all available information from the driver’s licence. * Scan a South African motor vehicle licence disc using a mobile device camera. * Read and decode the barcode or machine-readable information on the vehicle licence disc. * Extract all available vehicle and licence disc information. * Automatically identify whether the scanned document is a driver’s licence or vehicle licence disc. * Return all decoded information in JSON format. * Be developed in Python or include a Python-based backend component. * Be suitable for integration into our existing web-based software. * Work on Android and iOS mobile devices through a browser or suitable mobile interface. * Include validation and clear error messages when a document cannot be scanned or decoded. * Process all personal and vehicle information securely. * Avoid permanently storing information unless instructed by our software. * Be suitable for deployment on our own servers. ## Expected Driver’s Licence JSON Response The final JSON structure will depend on the information available on the licence but should include fields similar to the following: ```json { "document_type": "drivers_licence", "id_number": "", "surname": "", "initials": "", "date_of_birth": "", "gender": "", "licence_number": "", "licence_issue_number": "", "licence_issue_date": "", "licence_expiry_date": "", "licence_codes": [], "vehicle_restrictions": [], "driver_restrictions": [], "professional_driving_permit": { "code": "", "expiry_date": "" } } ``` ## Expected Vehicle Licence Disc JSON Response The final JSON structure will depend on the information available on the vehicle licence disc but should include fields similar to the following: ```json { "document_type": "vehicle_licence_disc", "registration_number": "", "licence_number": "", "vin": "", "engine_number": "", "make": "", "model": "", "colour": "", "vehicle_category": "", "vehicle_description": "", "licence_expiry_date": "", "disc_issue_date": "", "tare": "", "gross_vehicle_mass": "", "registering_authority": "" } ``` The developer must clearly identify which fields can reliably be extracted from the barcode and which fields may require OCR or another scanning method. ## Deliverables The successful freelancer must provide: 1. Complete Python source code. 2. A working South African driver’s licence decoder. 3. A working South African motor vehicle licence disc decoder. 4. Mobile camera scanning functionality or clear instructions for integrating mobile camera capture. 5. A REST API, Python library or reusable software component that can be incorporated into our existing system. 6. Structured JSON output for both document types. 7. Automatic document-type identification. 8. Installation and integration documentation. 9. A list of all libraries, dependencies and system requirements. 10. Sample integration code. 11. Error handling and data validation. 12. Testing using valid South African driver’s licence and vehicle licence disc samples. 13. Assistance during the initial integration and testing process. 14. Full ownership of the completed source code upon payment. ## Important Technical Requirements The solution should preferably: * Run on our own servers. * Not rely on a third-party decoding service. * Not require recurring licence or subscription fees. * Support mobile camera images captured under normal lighting conditions. * Handle rotated, blurred, partially obstructed or incorrectly positioned images where reasonably possible. * Use image-quality checks before attempting to decode a document. * Return clear error codes and messages when the image quality is insufficient. * Comply with secure development practices and South African POPIA requirements. Any paid libraries, proprietary software, external APIs or recurring costs must be clearly disclosed before the project is awarded. ## Required Experience Applicants should have experience with: * Python development * Barcode decoding * Optical character recognition * Image processing * OpenCV or similar technologies * REST API development * JSON * Mobile camera integration * Secure handling of personal and vehicle information Previous experience decoding South African driver’s licences, vehicle licence discs or identity documents will be highly advantageous. ## Information to Include in Your Proposal Please include: * Your relevant experience. * Examples of similar barcode, vehicle-document or identity-document projects. * Confirmation of whether you have previously worked with South African driver’s licences or vehicle licence discs. * Your proposed technical approach. * The libraries or technologies you intend to use. * Whether processing will take place on the mobile device or on the server. * Which information will be obtained through barcode decoding and which information will require OCR. * Any known technical limitations. * Your estimated completion timeframe. * Your fixed project price. * Details of any ongoing fees or external dependencies. We may request a paid proof of concept before awarding the complete project. The proof of concept should demonstrate successful decoding of at least one South African driver’s licence and one South African motor vehicle licence disc. All driver, vehicle and personal information provided during development and testing must be treated as confidential and may not be copied, retained, shared or used for any other purpose. Skills: Python, JSON, Image Processing, Mobile Development
Fixed budget:
250 - 750 USD
2 days ago
|
|||||
|
AI Background Removal System Development
Applied
|
~783 - 1,567 USD
|
2 days ago |
Client Rank
- not enough data
-
|
||
|
I need a Senior Computer Vision and Deep Learning Engineer to build a complete production-ready AI background removal system for AKPRINTHUB, with quality and speed close to remove.bg.
The system must accurately remove backgrounds from people, passport photos, products and general objects, including difficult hair, beard, fur and complex edges. This is not a basic rembg installation and no hidden third-party paid API should be used. The developer will be responsible for the complete project, including AI model selection/fine-tuning, Python FastAPI backend, GPU optimization, frontend integration with my existing PHP/JavaScript website, testing, bug fixing, deployment and production launch. Required technologies include Python, PyTorch, OpenCV, image segmentation, alpha matting, ONNX/TensorRT, CUDA, Docker and GPU deployment. Complete source code, trained model weights, training scripts, deployment files and documentation must be handed over. Payment will be milestone-based after quality and speed testing against a private dataset. Skills: PHP, JavaScript, Python, CUDA, Docker, OpenCV, Pytorch, Computer Vision, Deep Learning, FastAPI
Fixed budget:
75,000 - 150,000 INR
2 days ago
|
|||||
|
GNN Model Developer for Smart Contract Vulnerability Detection
Applied
|
$100
|
3 days ago |
Client Rank
- Risky
1 open job
Registered: Feb 3, 2026
17:54
1
|
||
|
We are seeking a skilled GNN model developer to enhance our smart contract vulnerability detection system. The ideal candidate will have experience in optimization and attention mechanisms, and be open to incorporating novel techniques. The role involves developing and refining GNN models to improve detection accuracy and efficiency. This is a part-time position for 1 to 3 months, requiring intermediate proficiency.
Fixed budget:
100 USD
3 days ago
|
|||||
|
Rust Desk
Applied
|
not specified | 4 days ago |
Client Rank
- Risky
1 open job
15:24
1
|
||
|
Branded Rust Desk client connected to my own self-hosted server (my host, my relay, my public key)
Budget:
not specified
4 days ago
|
|||||
|
Interactive Entertainment Robot Prototype
Applied
|
~8 - 13 USD
/ hr
|
4 days ago |
Client Rank
- not enough data
-
|
||
|
I’m ready to turn a playful idea into a working entertainment robot and need an engineer-maker who can take it from concept to functioning prototype. The goal is a small, crowd-pleasing bot that moves with personality, reacts to simple audience cues, and runs safely for several hours at events.
Scope • Mechanical design: compact chassis, expressive head/arm movement, easy-access battery housing. • Electronics: controller board selection (Arduino, Raspberry Pi, or your preferred MCU), motor drivers, Li-ion power system, safety cut-offs. • Firmware & control software: motion routines, basic sensor-based interaction (ultrasonic or camera vision for proximity), and a USB/Bluetooth interface for future updates. • Prototype assembly: sourcing parts, 3-D printing or light CNC where needed, complete bench testing, and a short demo video that proves mobility and interaction. Acceptance criteria 1. Robot rolls or walks smoothly on flat indoor surfaces. 2. Executes at least three distinct “show” motions triggered by a button press or distance sensor. 3. Operates 60 min minimum on a single charge without overheating. 4. All CAD files, schematics, BOM, and commented code delivered in editable formats. I’m open to your suggestions on components and toolchains—ROS, OpenCV, or custom libraries are all fine as long as setup instructions are included. Let’s build something people will remember. Skills: Electronics, Microcontroller, PCB Layout, Robotics, Arduino, 3D Printing, Mechanical Design, Firmware Development
Hourly rate:
750 - 1250 INR
4 days ago
|
|||||
|
AI Autism Behavior Analysis Webapp
Applied
|
~130 - 389 USD
|
5 days ago |
Client Rank
- not enough data
-
|
||
|
I am building a proof-of-concept web application that lets users upload an MP4 video and instantly receive an AI-driven assessment of autism-related behavioural cues. The core model will be a CNN-LSTM pipeline developed in Python with TensorFlow/Keras; OpenCV will handle video decoding and frame preparation, while MediaPipe will extract skeletal landmarks that feed into the network.
The model must recognise three categories of behaviour—repetitive movements, social interactions and emotional expressions—and return class probabilities for each. When a user presses “Analyse”, the site should process the clip server-side, run inference, then reunite the original footage with colour-coded video overlays that highlight the detected patterns frame by frame. The result should stream back to the browser inside a responsive HTML/CSS/JavaScript front end that feels clean on both desktop and mobile. What I will supply: • Any existing datasets, papers and partial code I have gathered so far. • Clarification on the CNN-LSTM architecture that is already showing promising offline results. What I need from you: • Finish or refactor the training script so it ingests MediaPipe keypoints and outputs saved model files ready for inference. • Build the Flask/FastAPI (or similar) back-end endpoint that accepts an MP4 upload, calls the model, and sends back processed frames or a video stream. • Create the front-end pages, progress bar, and final overlay visualiser. • Package everything so I can deploy it on a modest VPS or cloud instance with minimal setup. Acceptance criteria 1. User can upload an MP4 (≤200 MB), click Analyse and see a progress indicator. 2. After processing, the browser plays back the original video with overlay graphics marking detected repetitive motions, social cues and emotion cues in real time. 3. End-to-end latency for a 30-second clip on a mid-range GPU stays under two minutes. 4. Codebase is documented, uses virtual-env or Docker, and runs on Python 3.10+. If this matches your expertise in TensorFlow, MediaPipe, OpenCV and modern web stacks, let’s talk about timelines and milestones. Skills: JavaScript, Python, CSS, HTML5, OpenCV, Web Development, Flask, Keras
Fixed budget:
12,500 - 37,500 INR
5 days ago
|
|||||
Related freelance jobs queries: