Frequently Asked Questions About Private Dedicated AI
Everything you need to know about dedicated AI servers, data sovereignty, pricing, compliance, and deployment — answered plainly.
If your question isn't here, ask us directly through the contact form on the homepage.
No assumptions. No templated answers.
All Questions
What is a private AI server?
A private AI server is dedicated infrastructure that runs an AI model for a single organization. With Convergence AI, your own physical dedicated server hosts your model — no shared GPUs, no multi-tenant cloud, no third-party handoffs. Your data never touches a public AI provider's infrastructure.
How is a dedicated AI server different from cloud AI?
Cloud AI runs on shared, multi-tenant infrastructure where your data can be processed on hardware you don't control. A dedicated server is single-tenant: your data stays in your dedicated environment, with chain of custody from intake to inference.
What kind of organizations use Convergence AI?
Healthcare practices and hospitals, law firms, CPA firms, RIAs, insurers, mortgage lenders, and any organization handling sensitive or regulated data that can't afford to share infrastructure.
What are some examples of tasks a 400b server package could take over or streamline for a business?
This isn't just a chat window; it's an AI that's trained on your data, connected to your systems, and working alongside your team around the clock.
Knowledge management and internal operations. Your 400B model can ingest your entire knowledge base — employee handbooks, SOPs, past project documentation, training materials — and answer any question instantly. New hires get onboarded faster because they can ask “How do we process a vendor invoice?” and get a step-by-step answer pulled from your actual procedures. It can draft internal policies, summarize meeting transcripts, and flag inconsistencies between documents. No more digging through shared drives or asking colleagues who've been there for years.
Customer support and client communication. The model can be trained on your product catalog, pricing, and common support tickets. It can draft personalized responses to customer inquiries, escalate complex issues to human agents with full context attached, and even handle routine follow-ups like appointment reminders or order status updates. Because it's trained on your tone and your offerings, the responses don't sound like generic chatbot filler — they sound like your company.
Data analysis and reporting. If you connect it to your CRM, accounting system, or operational databases, the 400B can generate weekly sales summaries, flag anomalies in spending, identify trends in customer behavior, and draft board-ready reports. Ask it “What were our top three revenue drivers last quarter and where should we focus next?” and it will pull from your actual data, reason through it, and produce a clear, well-structured answer with supporting numbers.
Document review and contract work. For any business that handles contracts, vendor agreements, or legal documents, the model can review a 50-page agreement against your company's standard terms, flag deviations, summarize obligations, and draft suggested edits. It knows your risk thresholds because it was trained on your historical agreements. This alone can save legal or procurement teams dozens of hours per month.
Automated workflow execution. This is where the agentic advantage comes in. Your 400B can be given credentials to your web server, Slack, email, and other tools — and actually take over routine tasks. It can update your website, manage email lists, generate and post social content, monitor your server logs for issues, and even run data migration scripts. It executes multi-step processes end-to-end, not just answering questions about them.
Research and competitive intelligence. Give it access to industry reports, news feeds, and your internal strategy documents — it can track competitors, summarize market shifts, and prepare briefs for leadership. It's not browsing the open web like a general chatbot; it's synthesizing from sources you've authorized, with your business context in mind.
The through-line across all of these: the model already knows your business because it was trained on it. It's not starting from zero with every chat session. And because it's on your dedicated server with unlimited usage, your team can actually use it all day without watching token counts or worrying about costs.
Can the AI connect to our existing systems?
Yes. Direct integration into your existing internal systems — your team interacts with AI that already understands your business, your knowledge base, and your workflows. Not a bolted-on chat window.
Do I need to buy my own hardware?
No. Convergence AI builds and hosts your dedicated AI server in its secure private facility with direct connectivity to your network and team — the sovereignty of your own infrastructure without operating it yourself.
Does my data train public AI models?
No. On a dedicated Convergence AI server, every input, every output, and every model weight stays under your control. Your data stays in your dedicated environment and never trains a shared public model.
I was under the impression it's never safe to send passwords or API data through a chat session — sessions can leak to employees or third parties. Is that true with Convergence AI?
Your concern is valid — for SaaS-based shared AI platforms, whose computing happens in third-party data centers on multi-tenant infrastructure. It does not apply to Convergence AI. The entire purpose of a Convergence dedicated server is the opposite of shared: no shared tenancy, no shared computing, full custody of your private data. Your model runs on hardware dedicated to you alone. All communications travel over encrypted, secure connections and never leave your server. There is no other customer's session, data, or model anywhere near yours — and no public provider's infrastructure in the path.
How do I trust my data is secure?
Your data, your files, and your intellectual property never leave your building. Much like an employee accessing files on your network, your private AI — living on its own dedicated server — connects to that data on your terms. When your AI creates files, they are saved to the directory you specify on your network or web server.
How does your AI access, manage, and take over our accounts?
Deliberately simple — and deliberately under your control. Create a plain text file (like Notepad) named .env in a folder on your network or web server, enter any logins, credentials, or API keys you want your Convergence AI to manage, then point your AI to that file's location from your Convergence dashboard — and update it whenever you need. The file stays on your side, stored locally, never uploaded, and never leaving your control. Your dedicated AI reads it the same way an employee uses keys you gave them — and you can change or remove those keys at any time. It is really no different from handing a new employee the keys to your office doors.
What credentials do we provide our AI?
Whatever you want it to take over — no more, no less. Access to your network, if you want it learning from all your data. Access to Slack, Teams, Instagram, X, and similar platforms, if you want it working within or running those accounts. FTP or root access to your web server, if you want it to update, edit, and build out your website, handle email accounts, or perform server maintenance. You decide the scope as you go — start with read access if you prefer, watch it work, then expand. Every credential can be changed or revoked at any time.
What kind of firewalls or security does Convergence AI have to protect my dedicated AI server?
Your server is protected at every layer. Network: every Convergence server sits on its own isolated segment behind a dedicated enterprise firewall with default-deny policies — only the connections you approve are ever open. Management: administrative access runs on a separate private plane with multi-factor authentication, and every administrative action is logged and auditable. Data: encrypted in transit and at rest. Monitoring: our own AI watches the fleet 24/7 — unusual access patterns, traffic, or behavior trigger an immediate response. Physical: the facility is secured with badged and biometric access plus video surveillance, and power is protected by battery backup with redundant connectivity, including satellite failover. Maintenance: security patching is automatic and continuous. And because your deployment is single-tenant, there is no shared perimeter to compromise — your security boundary is your server, and nothing else lives inside it.
What compliance standards does it support?
Dedicated deployment is designed for regulated environments: healthcare (HIPAA), legal (privileged communications), and financial (confidential account data) — with full audit trails, access controls, and data that never leaves your controlled environment.
Is my data exportable if I leave?
Yes. Your data, your files, and your configurations are yours and fully exportable at any time. The Convergence model itself is licensed, not sold: when service ends, the model, its memories, and all data on the server are securely shredded, and the hardware is reprovisioned for the next client. Nothing of yours — and nothing of ours — carries over.
How did Convergence AI build its own AI / LLM? What LLM are you using?
Every dedicated server we deliver runs a proprietary Convergence AI model — distilled, sized, and trained for your industry and tooling. There is no off-the-shelf public chatbot behind your brand and no shared model behind your work. Sizes start at 30 billion parameters and scale to 400 billion and beyond for clients who need the top end. Because every client's requirements are unique, we establish yours during onboarding — and the toolbox your model carries grows as your needs grow. We continuously update and add the latest innovations, so your AI always has the best abilities at its disposal.
How is the model trained?
Your model is trained on your industry and your data — then deployed to your specifications. It continues to evolve with regular updates, retraining, and optimization as your business grows.
What does the 'b' mean in 30B to 400B?
B stands for billions of parameters. A parameter is a learned weight inside the neural network — think of it as a connection between concepts, built during training. More parameters means more capacity to store knowledge and reason. The rough rule of thumb: more parameters means better reasoning, deeper knowledge, and fewer mistakes — but it also requires more compute to run.
30B — fast and efficient for everyday work: summarization, drafting, data extraction, customer support responses, and classification. Ideal for high-volume routine tasks where you need speed and low overhead.
70B — a clear step up: complex reasoning, multi-step instructions, coding, and professional-grade drafting. For most business use cases — contract review, financial analysis, report generation, patient communication — this is the sweet spot.
400B — the high end: frontier-scale reasoning, nuanced understanding, and complex problem-solving across domains. Built for comprehensive legal research and drafting, multi-year financial statement analysis, complex medical cases, and sophisticated strategic work.
One more thing: size isn't everything — architecture and training matter too. A 30B model running on dedicated private infrastructure with your data will outperform any general-purpose SaaS chatbot, of any size, on your specific business tasks. A model tuned for your domain and used only by your people beats a bigger generalist every time.
Which model size should we choose?
30B suits solo practitioners and small teams; 70B suits small-to-medium firms handling sensitive multi-document work; 400B suits organizations where complex, high-stakes decisions demand the best AI available.
What hardware are you providing us?
Every Convergence server is engineered and specced at enterprise level in collaboration with NVIDIA — purpose-built for dedicated AI work, not repurposed desktop equipment. We also kept environmental impact front and center: low power draw by design (a fraction of hyperscale requirements), room-temperature operation at roughly 72 degrees Fahrenheit with no industrial chillers and no water-cooling infrastructure, and resilience by default — every server is monitored by our own AI 24/7, runs on battery backup, and connects through Starlink satellite internet, so short power or network outages never affect your operation.
How many tokens per second will our LLM do? What latency should we expect?
Tokens per second (tok/s) measures how fast the model generates text — roughly the reading speed of your AI. Latency is how quickly it starts answering. Both are measured on dedicated hardware, so results stay consistent around the clock, with no shared-queue slowdowns.
Tier 1 (16GB GPU, 30B model): roughly 60-100+ tokens per second, with responses starting in under a second. Feels instant — ideal for interactive chat, customer support, and high-volume routine work.
Tier 2 (32GB GPU, 70B model): roughly 30-50 tokens per second, with responses starting in about a second. Strong reasoning at a natural conversational pace — the sweet spot for most firms.
Tier 3 (128GB GPU, 400B model): roughly 15-35+ tokens per second, with responses starting in 1-2 seconds. Deep analysis takes a beat longer — the way it should.
Exact throughput depends on context length, tool use, and document size. During onboarding we measure your actual workloads and tune the server so the performance you care about — not benchmark numbers — is what you get.
How is pricing structured?
Capacity-based pricing: no per-use pricing, no token caps, no rate limits, and no surprise overage charges. Your server runs at your pace, on your schedule, with predictable cost based on your hardware specifications.
Where is the data center located?
Our hosting facility is in Boca Raton, Florida, with a second location opening in Dallas, Texas in the first quarter of 2027. Your dedicated server and your data remain onshore in the United States at all times.
Is the data center environmentally friendly?
Yes. Convergence servers are low-watt by design and room-temperature cooled — no industrial cooling infrastructure and no water waste. Hosting delivers AI with a fraction of the carbon footprint of hyperscale farms.
Still Have Questions? We'll Answer Them Straight.
We map your industry, workflows, data sources, and AI needs — no assumptions, no templated solutions. Tell us what you're protecting and we'll show you what dedicated AI looks like for your organization.