Just two years ago, “in-house AI” meant one thing for most companies: a ChatGPT or Copilot subscription. Today, more and more organisations are reaching the point where the cloud is no longer enough. There are usually three reasons: costs that rise with the number of queries, data that cannot be sent outside the company, and the need for AI to work with their own documents, databases and processes rather than general knowledge from the internet.

    This raises the question: what AI server should you actually buy, and how much does it cost?

    The answer is more interesting than it might seem, because the price range is huge: from a book-sized device costing just over twenty thousand złoty to a machine priced at more than two million. Crucially, each of these solutions makes sense, just for different users. Below, I outline three investment tiers, the calculation worth making before deciding, and five questions that will help you select the right hardware without buying too much or too little.

    How does an AI server differ from a standard server?

    A standard server is built around the CPU. An AI server is built around GPUs: they do all the work involved in running models, while the remaining hardware exists to support them. This results in three practical differences.

    • GPUs and their memory. Such a machine runs four to eight accelerators–most commonly as PCIe cards in rack servers, or as modules soldered onto a shared board in the most powerful platforms. Their total memory is the most important parameter in the entire purchase decision; more on that shortly.
    • Power. Eight cards under load consume 5–6 kW, equivalent to dozens of office computers. This requires an appropriately powered rack and efficient heat removal from the server room.
    • Networking. For a single machine, a standard high-speed network is sufficient. When several servers are combined into one system, much faster links are needed, because GPUs in separate chassis must exchange data with latency comparable to operation within a single chassis.

    This is why it is not simply a “more powerful server”. It is a different hardware category, selected according to different criteria.

    GPU memory: the one number that puts the whole selection into perspective

    There is one parameter more important than price: the total GPU memory in the server.

    Why? The entire model must fit into that memory, much like a file must be loaded into an application before it can be used. If it does not fit, it simply will not start.

    However, the model itself is only the beginning of the calculation. Every logged-in user and every document being analysed requires additional memory for the active conversation. The same model therefore needs noticeably more space for one hundred people than for one–and it is usually the number of users, rather than the size of the model itself, that determines which hardware tier is required.

    The guide covers the baseline memory requirements of models: Which AI computer should you choose?. The table below shows the second dimension of this calculation: how requirements change with the scale of deployment. The values are indicative and apply to a typical production configuration.

    Model size5 users50 users200 usersHardware tier
    small (8B class)approx. 8 GBapprox. 20 GBapprox. 50 GBDGX Spark
    medium (70B class)approx. 45 GBapprox. 70 GBapprox. 130 GBrack server, 1–2 cards
    large (200B+ class)approx. 130 GBapprox. 200 GBapprox. 350 GBeight-card server
    training a custom modelseveral times more than inference aloneHGX B200 platform

    Labels such as 8B and 70B indicate how many billions of parameters a model contains–8B means eight billion. The more parameters there are, the better the response quality and the more GPU memory the model requires.

    There is also one thing that is easy to overlook.

    Headroom. The figures above reflect normal operation. At peak load, it is worth adding a 20–30% margin; otherwise, the system will start rejecting queries when traffic increases.

    Tier 1: NVIDIA DGX Spark–the starting point

    The NVIDIA DGX Spark is a book-sized desktop device with 128 GB of unified memory and AI computing performance of around 1 petaFLOPS. It enables local deployment of models in the class that powers popular assistants, without sending a single document to the cloud. Several variants are available, differing mainly in storage capacity (1, 2 or 4 TB) and support package. The full range in this hardware class, including models from eight manufacturers, is available in the AI computers and workstations category.

    Who is it for? For a company that first wants to verify whether AI can improve customer service, contract analysis or quotation preparation before anyone commits a substantial budget. The investment is around £4,728.50 net–less than an annual subscription to AI tools for a team of more than a dozen people. Two devices can be combined into a single system.

    Limitation. This hardware is for a single team, not the entire organisation. It cannot be installed in a rack, has no redundant power supply and cannot be covered by the service policy used in the server room. When one hundred people need to use AI simultaneously, it is time to move up to the next tier.

    Tier 2: rack-mounted GPU server–production AI for the enterprise

    In my view, this is the tier most often overlooked today, yet it is the most sensible choice for companies with real AI use cases.

    This refers to a dedicated rack-mounted GPU platform (5U), designed for eight double-width cards. It is not tied to a single accelerator model–the choice of cards depends on the use case. The base configuration includes two AMD EPYC processors, up to 6 TB of DDR5 memory across 24 slots, and six 2700 W power supplies operating redundantly.

    The number of cards is key, as it determines the available GPU memory. Two RTX PRO 6000 cards with 96 GB each provide 192 GB and support an internal assistant for several dozen users. A full eight-card configuration delivers 768 GB, enough for several hundred users or very large models. Additional cards can be added in the same chassis without replacing the server – the biggest advantage at this level is that you do not need to invest in the full configuration from the outset.

    Specific model: Supermicro AS-5126GS-TNRT Gold Series

    Sold as a complete, configured system. Expected investment for the entry configuration with two cards: approximately £98,510.52 net. Check the product page for the exact specification and current price.

    What does this mean in practice? This type of server can run large language models locally, analyse a corporate document archive, automate customer service, generate content in multiple languages and analyse images from a production line. RTX PRO cards are also versatile: the same server can be used for rendering, engineering simulations or virtual workstation infrastructure. This is often a deciding factor: the hardware does not sit idle between AI projects.

    Who is it for? In the entry configuration, it is suited to organisations with several dozen employees that have completed the testing phase and want to give teams access to an internal assistant familiar with company procedures. In the full configuration, it is intended for large organisations, hundreds of concurrent users and regulated industries such as finance, healthcare and the public sector, where data must remain on site.

    What to keep in mind. Power consumption increases with the number of cards: from approximately 2 kW with two cards to 5–6 kW with eight. The rack must be able to supply this power, and the cooling system must be able to remove the heat. If you plan to expand, size the power and air-conditioning capacity for the target configuration from the outset, rather than the entry configuration.

    Level 3: NVIDIA HGX B200, the infrastructure used to train models

    At the very top end is the type of hardware featured in reports on the technology race between the world’s largest companies. A server based on the NVIDIA HGX B200 platform includes eight Blackwell GPUs, currently the most powerful AI accelerators on the market, with a total of 1,44 TB of HBM3e memory – the fastest currently available.

    The machine occupies 10U in a rack and has eight 400 Gb/s network interfaces, as it is designed to be clustered with additional identical servers. Power is supplied by six power supplies rated at approximately 5 kW, operating redundantly in a 3+3 configuration – actual system consumption under load is around 10 kW, while the remainder provides capacity headroom and failure protection. It includes a BlueField-3 processor, which offloads networking and security tasks from the main system, as well as three years of next-business-day on-site support.

    Specific model: Supermicro AS-A126GS-TNBR

    Designed for scientific computing, AI, industrial automation and data analytics. Expected investment: approximately £394,042.08 net per server; see the product page for the current price.

    The difference compared with a Level Two GPU server is not that B200 “does the same thing, only faster”. It performs tasks that the other system cannot complete in a reasonable timeframe: training proprietary models from scratch, fine-tuning models with hundreds of billions of parameters, supporting hundreds of users simultaneously, scientific research and industrial-scale simulations.

    Who is it for? For companies building AI-based products that create their own models rather than relying on ready-made ones. For research institutions and universities. For telecommunications operators, banks, large industrial groups and pharmaceutical companies.

    In practice, a single machine of this type is rarely purchased on its own. It also requires network infrastructure, power and data centre cooling – at 10 kW per chassis, liquid cooling is typically required. This is not simply a hardware purchase; it is a strategic decision.

    How do you choose an AI server? Five questions worth asking

    Do you already know what you want AI to do in your organisation?

    If not, choose Level One. For the price of one premium laptop, you get an environment in which you can test several ideas within two weeks and make decisions based on facts rather than presentations.

    How many people will use it, and how often?

    If the answer is “the whole company, every day”, a desktop device will not be enough, and a rack-mounted GPU server will be the right choice. It is worth calculating how much you currently spend on cloud APIs and subscriptions. At scale, owning your own hardware is more cost-effective over a two- to three-year period.

    Are you building your own models or using ready-made ones?

    Using ready-made models, even very large ones, is the domain of Level Two. Building proprietary models requires B200. The dividing line here is clear, and there is rarely any doubt about which side an organisation falls on.

    How large a model do you need to run, and for how many users?

    Refer back to the memory table. If your chosen model requires 60 GB and fifty people need to use it simultaneously, a single card with 96 GB of memory is the absolute minimum, while two provide headroom and room for growth.

    What can your data centre support?

    The most common surprise in AI projects is not the hardware but the building: how much power can be supplied, whether the air conditioning can remove the additional heat and whether there is space in the rack. It is better to check this before placing the order than after delivery.

    What else is included in the implementation budget?

    The server price typically accounts for 70–85% of the project budget. The remaining items should be included at the planning stage.

    • Power and rack infrastructure – A suitably rated PDU, phase distribution, uninterruptible power supply and, for some installations, a new electrical connection.
    • Cooling – up to approximately 5–6 kW per rack, precision air conditioning is usually sufficient. Above 10 kW, liquid cooling is used, which requires room adaptation and a separate service procedure. If the server room was not designed for this, it is worth including this cost in the quotation from the outset.
    • RAM – AI configurations start at several hundred gigabytes, and prices for server memory modules rose significantly in 2026. The difference between 512 GB and 1 TB can be a substantial cost item.
    • Networking – a switch with adequate bandwidth, transceivers and cabling. Cluster configurations also require a separate network for GPUs.
    • Storage – models and document collections occupy terabytes, and read throughput should match the performance of the cards.
    • Deployment – software installation, model deployment, integration with your company login system and the systems you use.
    • Support and warranty – for production environments, next-business-day on-site service is the standard.

    Where to start

    The most common mistake in purchases like this is starting with a catalogue. The sequence that works is the reverse: first, a single use case; then, the number of users; then, a comparison with current cloud costs; and only finally, the configuration.

    If you do not yet know exactly what AI will be used for, level one is sufficient for several months of testing and costs as much as one laptop. If testing is already behind you and the assistant is to serve the whole company, level two is currently the most sensible choice for most organisations. Level three is chosen when you are building your own models – and in that case, the decision is usually clear.

    If the configuration needs to be tailored to your server room – including the number of cards, memory, networking or cooling – our team will prepare a proposal for your specific use case.

    FAQ

    How much does an AI server cost?

    The range is wide: from approximately £4,728.49 net for a DGX Spark-class desktop device, through approximately £98,510.52 for an entry-level rack-mounted GPU server, to approximately £394,042.87 for an HGX B200 platform. Add 15–30% to the hardware price for power, cooling, networking and deployment.

    Is an AI server more cost-effective than the cloud?

    It depends on the scale. Cloud costs increase with the number of requests, while the cost of your own server is fixed – after purchase, you do not pay for every call. It is therefore worth comparing current API bills and subscription costs with hardware instalments spread over three years and energy costs. For intensive use, your own infrastructure comes out ahead; for occasional use, the cloud remains cheaper.

    How much GPU memory does a language model need?

    The entire model must fit into the memory of the cards. As a guide: a 8B-class model requires several dozen gigabytes; a 70B-class model requires from 70 to 140 GB, depending on the model precision; and models with more than 200 billion parameters start at several hundred gigabytes. In addition, memory is required for active conversations, increasing with the number of users.

    Is a standard server with a graphics card sufficient for AI?

    For testing and individual use cases – yes. For production workloads, usually not, because the limitations are the total memory across the cards, as well as the power delivery and cooling of a chassis not designed for eight accelerators.

    Which AI server should a company with up to 50 employees choose?

    If AI is to run in production for everyone, the most sensible option is a rack-mounted GPU server configured with two cards, with the option to add more. If you are still at the testing stage or it is used by a single team, a DGX Spark-class desktop device is sufficient, or alternatively two units connected as one system.

    How much electricity does an AI server use?

    A desktop device uses approximately 240 W. A rack-mounted GPU server uses from approximately 2 kW with two cards to 5–6 kW with eight, while an HGX B200 platform is around 10 kW. Calculate the cost based on your own kWh rate and planned operating time – in continuous operation, the difference between the levels is several dozen times over.

    Udostępnij.
    łukasz bojar

    Vice President Sales & MarketingHe has been with Senetic since 2014, where he is responsible for shaping the company’s vision and sales strategy. His professional ambition is to expand into new sales markets around the world. He has over 10 years of experience in management, business development and sales, gained both within and outside the IT industry. He shares his expertise as a speaker at industry conferences. In his free time, he relieves stress at the gym or by running in the park while discovering new genres of music.

    Dodaj komentarz