Sharepoint consultant(Only Locals)
Real Soft Inc.
Philadelphia, PA, United States
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Compensation
$104,000.0 - $124,800.0
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
C++ (Programming Language)
Python (Programming Language)
Metadata
Open Source Technology
Performance Tuning
Microsoft SharePoint
Data Logging
Large Language Models
Generative AI
Backend
Optimization Algorithms
+2 more
HuggingFace
Docker
Job description
- Deploy and manage open-source LLMs (e.g., Llama 3, Mistral/Mixtral) in on-prem or private environments
- Develop and optimize LLM inference workflows using Python
- Implement Retrieval-Augmented Generation (RAG) pipelines
- Design and integrate vector database solutions for efficient semantic search
- Perform model quantization and performance tuning for CPU-based inference
- Ensure data privacy, security, and governance compliance in enterprise environments
- Implement access controls, logging, and monitoring mechanisms
- Deliver reference architecture, prototypes, and technical documentation
- Collaborate with internal teams for knowledge transfer and system adoption
Requirements
We are seeking a mid-level Software Developer/Engineer with hands-on experience in deploying on-premise LLM solutions and vector databases. The ideal candidate will have strong expertise in Python, RAG pipelines, and enterprise-grade AI system implementation within secure environments., * Strong experience with Python for AI/ML and backend development
- Hands-on experience with open-source LLM deployment (Llama 3, Mistral, Mixtral)
- Experience with CPU-based inference and optimization techniques
- Practical experience with vector databases (Qdrant, Chroma, Milvus, pgvector)
- Proven experience building RAG pipelines
- Knowledge of embeddings, similarity search, and metadata filtering
- Understanding of enterprise security, data privacy, and air-gapped environments
Preferred Qualifications
- Experience with LangChain or LlamaIndex
- Familiarity with Docker and Kubernetes
- Exposure to Rust, Go, or C++ for high-performance systems
- Experience with LLM inference frameworks (vLLM, llama.cpp, Hugging Face Transformers)
- Prior experience in regulated or enterprise environments
Deliverables
- End-to-end reference architecture for LLM + vector DB solutions
- Functional prototype (LLM + RAG + Vector DB)
- Comprehensive documentation and knowledge transfer
Benefits & conditions
4.04.0 out of 5 stars Philadelphia, PA 19122 Hybrid work $50 - $60 an hour - Contract
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
over 2 years ago
CH
Chris Heilmann
Dev Digest 132 - Binging WADFlix?
almost 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
CH
Chris Heilmann
Dev Digest 138 - Are you secure about this?
almost 2 years ago