Introduction
DanteGPU is a GPU-as-a-Service (GPUaaS) project designed to democratize access to high-performance computing. At its core, DanteGPU enables individuals and entities to monetize their unused graphics card capabilities by renting them to a distributed network. Users of the platform can earn $dGPU tokens for providing these resources.
The project features Agora, an all-in-one Agent marketplace, and emphasizes AI interoperability. This allows users to access and utilize various AI Agents on a pay-per-use basis, billed hourly, thereby avoiding hefty monthly subscription costs. Furthermore, DanteGPU empowers users to publish their own AI agents on the Agora marketplace, fostering a rich ecosystem of AI tools and services.
Core Mission: DanteGPU aims to dismantle the centralized control currently held by a few major players in the AI and high-performance computing landscape. By eliminating central points of control, the project strives to make powerful GPU resources accessible to everyone, from individual AI hobbyists and researchers to small development teams. https://app.gitbook.com/o/G99KZfPVvkpFW0qw5X5t/s/7ZpP0HETQD2EI5hgSjvt/~/changes/10/dantegpu-backend/core-services
Key Tenets of DanteGPU:
-
Democratized GPU Access:GPU providers can directly offer their unused resources to AI developers and researchers, ensuring fair and efficient access to computing power without intermediaries. -
Real-Time Access to Distributed GPU Resources:AI agents can instantly discover available GPU resources through a blockchain-powered marketplace (Agora) and scale their workloads efficiently. -
Autonomous GPU Selection & Optimization:AI agents are empowered to select GPUs based on specific criteria such as VRAM, processing power, and bandwidth requirements, while also optimizing for cost and performance. -
Flexible Usage Model & Lower Costs:DanteGPU champions a pay-as-you-go model, removing the need for long-term subscription commitments typically found in traditional cloud GPU services. -
Accelerated Model Training:The distributed nature of DanteGPU’s resources allows for the parallelization of AI model training workloads across multiple GPUs, significantly speeding up development cycles. -
Secure & Transparent Transactions:Leveraging the Solana blockchain, DanteGPU ensures that all transactions between GPU providers and AI agents are transparent, immutable, and securely recorded.
This documentation will guide you through the various aspects of the DanteGPU project, including its architecture, components, and how to participate as a user or provider.
Project Vision, Scope, and Target Audience
This section provides a deeper insight into the fundamental goals of the DanteGPU platform, the specific challenges it addresses, the key parties involved in its development, and its intended users, with a clear distinction for its AI marketplace product, Agora.
Project Name and Core Concept
DanteGPU is a comprehensive platform built on the Solana blockchain, designed to democratize access to GPU resources and AI capabilities. A key product within the DanteGPU ecosystem is Agora, an open marketplace for AI agents, chatbots, and AI-enhanced applications. The platform as a whole aims to create a environment where providers can offer GPU power and AI tools, and consumers can access them using the unified utility token, $dGPU.
Primary Business Problem Addressed
The current AI landscape is predominantly controlled by a handful of large corporations possessing extensive computational resources. This centralization creates barriers to entry and innovation for smaller entities and individuals. The DanteGPU platform, through its products and services including the Agora Marketplace, aims to address this critical business pain point by:
-
Democratizing Access to AI Tools: Allowing anyone, from individual developers to research institutions, to publish their AI-enhanced tools (AI-powered applications, chatbots, image generators, agents, etc.) on
Agora, DanteGPU’s open marketplace. -
Democratizing Access to GPU Power: Enabling individuals and entities to rent out their unused GPU capacity to the network, making high-performance computing more accessible.
-
Unifying Utility: Enabling users to access a diverse range of AI-enhanced tools on
Agoraand GPU resources using a single, central utility token, $dGPU, simplifying transactions and fostering a cohesive ecosystem.
By tackling these issues, DanteGPU seeks to foster a more equitable, innovative, and accessible AI ecosystem.
Key Stakeholders
While originating as a hackathon project that did not secure grant funding, the DanteGPU platform is under continuous and active development by its original founding team. This team remains the primary group invested in the project’s success, driving its vision, development, and future growth. As the platform evolves, other stakeholders will include:
-
GPU Providers: Individuals and entities who rent out their GPU resources to the DanteGPU network.
-
AI Tool Publishers/Providers on
Agora: Developers and organizations who list their AI agents and applications on theAgora Marketplace. -
Consumers/End-Users: Individuals and businesses who utilize the AI tools on
Agoraand GPU resources available through the DanteGPU platform. -
$dGPU Token Holders: Individuals invested in the ecosystem’s utility and growth.
Project Timeline and Constraints (Current Status)
The initial roadmap for the DanteGPU platform includes the target of a beta launch for the Agora marketplace approximately seven month from the project’s effective restart (post-hackathon development phase), followed by a public release after an additional month of rigorous testing and refinement.
Current known challenges and constraints for the DanteGPU platform include:
-
Expertise Gaps: Limitations in specialized knowledge concerning system architecture for high-availability systems.
-
Scalability Engineering: Designing and maintaining web infrastructure capable of handling high traffic volumes, particularly for
Agora. -
Blockchain Operations: Managing a live blockchain product that involves real-time payment processing for both GPU rental and
Agora Marketplacetransactions, and the handling of sensitive user and transaction data.
These areas are actively being addressed by the development team.
Expected Business Outcomes and Success Metrics
The success of the DanteGPU platform will be measured by a combination of platform adoption, ecosystem growth (for both GPU sharing and Agora), and token utility. Key metrics include:
-
User Acquisition: The total number of registered users on the platform (both consumers and providers of GPU resources and AI tools).
-
Marketplace Activity (
Agora): The number of unique AI-enhanced applications and agents published onAgora. -
GPU Resource Utilization: The amount of GPU compute time rented through the platform.
-
Provider Earnings: The cumulative amount of $dGPU earned by AI tool creators on
Agoraand GPU providers. -
Token Velocity & Valuation: The indirect increase in $dGPU’s market capitalization (MCAP) as a reflection of platform utility, driven by consumers exchanging $dGPU for services (potentially via $SOL or other pairs) and creators/providers being rewarded in $dGPU.
Additional relevant metrics may be identified and incorporated as the platform matures.
Target Audiences
The DanteGPU platform is designed to serve two primary groups, particularly concerning its Agora marketplace and GPU rental services:
For AI Tool Creators/Publishers (on Agora):
The ideal publisher for the Agora Marketplace includes, but is not limited to:
-
Individual AI Hobbyists: Enthusiasts developing novel or niche AI tools.
-
Small AI Development Teams: Startups and independent teams creating specialized AI solutions.
-
Academic Researchers: Institutions and researchers looking to share their AI models and tools with a broader audience or monetize their research via
Agora. -
Developers of Foundational Models: Those who wish to provide access to large, pre-trained language models that can serve as a base for various specialized tasks through
Agora.
For GPU Providers (on DanteGPU):
-
Individuals with powerful gaming PCs or mining rigs with idle GPU capacity.
-
Small to medium-sized data centers or businesses with underutilized GPU servers.
-
Researchers or institutions with specialized GPU hardware available for specific tasks.
For End-Users (Consumers on DanteGPU & Agora):
The primary target end-users for the AI tools available on Agora and the GPU resources on the DanteGPU platform are:
-
Businesses: Companies of all sizes seeking to integrate AI capabilities (via
Agora) or access raw GPU power for model training/inference without significant upfront investment in hardware or complex monthly subscriptions. -
Individual Consumers: Users looking for specific AI tools on
Agorafor personal use, creative endeavors, or learning purposes, or individuals needing temporary GPU power. -
Developers: Programmers and software engineers who wish to integrate AI functionalities from
Agorainto their own applications or require GPU resources for their development projects. -
AI Developers and Researchers: Those needing access to diverse and scalable GPU resources for training complex models or running intensive simulations.
Architectural Overview
The DanteGPU platform is engineered as a distributed system, with its backend services forming the operational core responsible for orchestrating GPU resources, managing AI workloads, ensuring secure access, and facilitating a seamless user experience within the Agora marketplace and the broader GPU-as-a-Service (GPUaaS) ecosystem. The architecture is predicated on modern principles of microservices, robust communication patterns, and a carefully selected technology stack to achieve high availability, scalability, and maintainability.
I. Architectural Paradigm: Decentralized Microservices
The DanteGPU backend eschews a monolithic structure in favor of a microservices architecture. This paradigm involves decomposing the application into a suite of small, independently deployable services, each organized around a specific business capability or domain. This choice is foundational to achieving the platform’s goals of democratization and distributed operation.
A. Rationale and Benefits:
-
Enhanced Scalability & Elasticity: Each microservice can be scaled independently (e.g., horizontally by adding more instances or vertically by allocating more resources) based on its specific load profile. For instance, the
Job Queue Servicemight require different scaling characteristics during peak submission times than theAuthentication Service. This granular scalability ensures optimal resource utilization and cost-effectiveness. -
Improved Fault Isolation & Resilience: The failure of a single microservice, if designed correctly with patterns like circuit breakers or bulkheads (though specific implementations depend on service logic), is less likely to cause a catastrophic failure of the entire platform. This resilience is paramount for a system managing real-time GPU access and financial transactions ($dGPU).
-
Technology Diversity & Specialization: While the primary languages are Go and Python (FastAPI), a microservices approach theoretically allows for selecting the best technology (language, database, etc.) for each service’s specific needs without impacting others. This fosters innovation and allows developers to leverage optimal tools for particular problems. For example, Go’s concurrency primitives and performance are ideal for network-intensive services like the API Gateway or resource orchestration, while Python’s rich ecosystem and rapid development capabilities suit services like user authentication or business logic layers.
-
Independent Development & Deployment Cycles: Teams can develop, test, and deploy their respective microservices autonomously. This accelerates development velocity, simplifies continuous integration/continuous deployment (CI/CD) pipelines, and reduces the scope and risk of individual deployments.
-
Clearer Domain Boundaries (Domain-Driven Design - DDD): Microservices naturally align with DDD principles, where service boundaries are defined around specific business domains (e.g., User Management, GPU Resource Management, Job Lifecycle Management). This leads to services with high cohesion and loose coupling, making the system easier to understand, evolve, and maintain.
-
Alignment with Decentralization Ethos: A distributed network of services mirrors the decentralized nature of the GPU providers and consumers DanteGPU aims to connect.
B. Inherent Challenges and Mitigation Strategies:
Adopting microservices also introduces complexities that the DanteGPU architecture must address:
-
Operational Overhead: Managing a multitude of services requires robust automation for deployment, scaling, monitoring, and logging. Technologies like Docker, Docker Compose, and a future transition to Kubernetes are key to mitigating this.
-
Distributed System Complexity: Debugging and tracing requests across multiple service boundaries can be challenging. Implementing distributed tracing (e.g., using OpenTelemetry) and comprehensive, correlated logging are essential.
-
Inter-Service Communication: Network latency, reliability, and the need for robust communication patterns (discussed below) become critical concerns.
-
Data Consistency: Maintaining data consistency across services that own their respective databases requires careful design, often employing patterns like eventual consistency, sagas, or two-phase commits where strong consistency is indispensable.
-
Testing Complexity: End-to-end testing of workflows spanning multiple services requires more sophisticated strategies than testing a monolith.
II. Inter-Service Communication Strategy
Effective communication between microservices is vital. DanteGPU employs a hybrid approach, leveraging both synchronous and asynchronous patterns:
A. Synchronous Communication:
Used for request/response interactions where an immediate response is expected.
gRPC (Google Remote Procedure Call):
-
Rationale: Preferred for internal, high-throughput, low-latency communication between backend services.
-
Mechanism: Utilizes HTTP/2 for transport, offering multiplexing, header compression, and bidirectional streaming. Protocol Buffers (Protobufs) are used as the Interface Definition Language (IDL), enforcing contract-first design, ensuring type safety, and enabling efficient binary serialization/deserialization.
-
Benefits: High performance, efficient data encoding, strongly-typed contracts, support for streaming, and code generation in multiple languages.
-
Use Cases: Internal API calls between core services like the
Scheduler/Orchestrator Servicequerying theProvider Registry Servicefor available GPUs, or internal control plane operations.
RESTful APIs (HTTP/JSON):
-
Rationale: Employed for services that might be consumed by a wider range of clients (including potentially third-party developers in the future) or where the overhead of gRPC setup is not justified. The API Gateway also exposes RESTful endpoints to external clients.
-
Mechanism: Standard HTTP methods (GET, POST, PUT, DELETE) with JSON payloads. Adherence to REST principles (statelessness, resource-based URLs) is expected.
-
Benefits: Simplicity, ubiquity, human-readability (JSON), wide support across languages and tools, easier integration with web frontends.
-
Use Cases: External API Gateway endpoints, specific internal services where simplicity is prioritized over raw performance, or interaction with services like the
Authentication Servicefrom the API Gateway.
B. Asynchronous Communication / Event-Driven Architecture:
Used for decoupling services, improving resilience, and handling long-running or background tasks. This is crucial for a system managing potentially time-consuming AI jobs.
- NATS JetStream:
-
Rationale: Provides a persistent, reliable, and high-performance messaging and streaming platform for asynchronous operations.
-
Mechanism: NATS is a lightweight, high-performance messaging system. JetStream adds persistence, message replay, and various delivery semantics (at-least-once, at-most-once, and potentially exactly-once patterns depending on consumer logic). Services publish events/messages to named “subjects” (topics), and interested services subscribe to these subjects.
-
Streams: Persistent logs of messages.
-
Consumers: Allow services to read messages from streams, with options for push or pull delivery, acknowledgments, and durable subscriptions.
-
-
Benefits:
-
Decoupling: Producers and consumers are independent; they don’t need to know about each other or be available simultaneously.
-
Resilience & Durability: Message persistence ensures that requests are not lost if a consuming service is temporarily unavailable.
-
Scalability: Allows for scaling consumer groups independently to process messages in parallel.
-
Load Leveling: Smooths out peak loads by queuing requests.
-
-
Use Cases:
-
Job Queuing: The
Job Queue Servicerelies on NATS JetStream to persist AI job requests submitted by users. TheScheduler/Orchestrator Serviceconsumes these jobs from the queue. -
Event Notification: Broadcasting events like “GPU available,” “job status updated,” or “new model published on Agora” to interested services without direct coupling.
-
Data Pipelines: Facilitating asynchronous data flows, e.g., logs or metrics forwarding before final aggregation.
-
III. Core Technology Stack Choices
The selection of technologies for the DanteGPU backend reflects a pragmatic approach, balancing performance, developer productivity, and ecosystem support.
-
Go (Golang):
-
Role: Primary language for high-performance, concurrent network services and infrastructure components.
-
Strengths: Excellent support for concurrency (goroutines, channels), compiled to native code for speed, static typing for reliability, efficient memory management, and a strong standard library for networking. Ideal for services like the
API Gateway (Siger),Provider Registry Service, andScheduler/Orchestrator Service.
-
-
Python (with FastAPI):
-
Role: Used for services where rapid development, a rich ecosystem of libraries (e.g., for machine learning, data science, web frameworks), or specific integrations are key.
-
Strengths (FastAPI): Modern, high-performance web framework for building APIs with Python 3.7+ based on standard Python type hints. Offers automatic data validation, serialization, interactive API documentation (Swagger UI, ReDoc), and leverages Starlette (for web parts) and Pydantic (for data parts). Ideal for the
Authentication Service.
-
-
Docker & Docker Compose:
-
Role: Containerization technology for packaging applications and their dependencies. Docker Compose is used for defining and running multi-container Docker applications, especially in development and testing environments.
-
Benefits: Environment consistency, isolation, portability across machines, simplified dependency management, and a foundational step towards more advanced orchestration.
-
-
Consul (by HashiCorp):
-
Role: Service discovery, configuration management, and health checking.
-
Mechanism: Services register themselves with Consul, and other services can query Consul to find their network locations. Consul performs health checks to ensure only healthy service instances receive traffic. It can also serve as a distributed key-value store for dynamic configuration.
-
Benefits: Enables dynamic scaling and resilience, as services don’t need hardcoded addresses. Simplifies the routing logic in the API Gateway and internal service communication.
-
-
PostgreSQL:
-
Role: Robust, open-source object-relational database system used for persistent storage by services requiring structured data and transactional integrity.
-
Strengths: ACID compliance, reliability, extensibility, rich feature set (JSONB support, full-text search, etc.). Suitable for the
Provider Registry Service,Scheduler/Orchestrator Service(job store), andAuthentication Service(user data).
-
-
MinIO:
-
Role: High-performance, S3-compatible object storage service.
-
Strengths: Scalable, resilient storage for unstructured data like AI models, datasets, job results, and user uploads. Can be self-hosted, providing data sovereignty.
-
-
Kubernetes (Future Consideration):
-
Role: Advanced container orchestration platform.
-
Benefits: Automated deployment, scaling, self-healing, service discovery, load balancing, and configuration management for containerized applications at scale. A natural evolution from Docker Compose for production environments demanding higher resilience and operational efficiency.
-
IV. Overview of Core Service Domains
The backend is logically segmented into several domains, each encompassing one or more microservices:
-
API Gateway (
siger-api-gateway): The unified ingress point. Handles routing, authentication (JWT), rate limiting, CORS, and acts as a facade for the backend services. (To be detailed in Page 5). -
Authentication Service (
auth-service): Manages user identities (providers, consumers), registration, credential verification (password hashing), JWT issuance, and potentially profile management. -
Provider Registry Service (
provider-registry-service): Tracks connected GPU providers, their hardware specifications (GPU model, VRAM, drivers), real-time status (idle, busy), location, and utilization metrics. Critical for the scheduler to find suitable resources. -
Job Queue Service (Integrated via NATS): Manages the intake and persistent queuing of AI job requests.
-
Scheduler/Orchestrator Service (
scheduler-orchestrator-service): The “brain” of the system. Dequeues jobs, queries the Provider Registry for suitable GPUs, dispatches tasks (likely via NATS) to provider daemons (beatrice-core-services), and tracks job progress. -
Storage Service (
storage-service): Abstracts interactions with object storage (MinIO), handling uploads, downloads, and management of models, datasets, and results. -
Monitoring & Logging Service (
monitoring-logging-service): A stack (Prometheus, Grafana, Loki, Promtail, etc.) for aggregating metrics and logs from all services, enabling observability and debugging. -
GPU Provider Daemon (
beatrice-core-services): A client-side agent running on the GPU provider’s machine. Responsible for receiving tasks from the Scheduler, executing them (e.g., in a containerized environment with GPU passthrough), monitoring execution, and reporting back results and status. (This component is critical for the GPUaaS functionality). -
Billing & Payment Service (Planned Post-MVP): Will integrate with payment gateways to track resource usage (GPU time, storage) and manage financial transactions (payouts to providers, charges to consumers).
V. Cross-Cutting Concerns
Several concerns span multiple services:
-
Security: Beyond authentication/authorization, includes secure inter-service communication (e.g., mTLS), input validation, protection against common vulnerabilities, and secure secrets management.
-
Observability:
-
Logging: Consistent, structured logging across all services, often correlated with request IDs for tracing.
-
Metrics: Collection of key performance indicators (KPIs) from each service for monitoring health and performance.
-
Distributed Tracing: Implementing mechanisms (e.g., OpenTelemetry) to trace requests as they flow through multiple services.
-
-
Configuration Management: Centralized and dynamic configuration for services, potentially using Consul’s KV store or environment variables managed by the orchestration platform.
-
Error Handling & Resilience: Consistent error reporting, retry mechanisms, and fault tolerance patterns within and between services.
For more information:
X -> https://x.com/dantegpuaas
Github Organization (Star Us) → DanteGPU · GitHub